← AI & LLM Work
In production Built with n8n

Cut AI Costs Without Rebuilding the Stack

A self-hosted automation layer routes every request to the cheapest model that can actually handle it — open-source and local where possible, a commercial or frontier model only where the task needs it. Nothing else in the stack changes; the routing decision happens before the request ever reaches a model.

One Request, Routed to the Right Model

Complexity decides the destination — not a fixed default.

Step 1

Request Comes In

Any channel, any volume — the same entry point regardless of source.

Step 2

Complexity Classified

A routing layer scores the request and picks the cheapest tier able to handle it.

Step 3

Served by the Right Tier

Self-hosted, mid-tier, or frontier — usage, cost, and latency are logged either way.

Measured in production, first month

Where the Traffic Actually Goes

~30%
low-complexity, served by self-hosted open-source models
40–50%
mid-complexity, served by mid-tier commercial models
~20–30%
reaches a frontier model — reserved for what needs it

Every re-routed request is a request that stops paying frontier per-token pricing. Defaulting everything to the most capable model was the baseline this replaced.

What the Automation Layer Gives You

The platform capability, independent of any one workflow.

Add Models in Minutes

Plug in a new open-source or hosted model without redeploying the platform — point the router at it and traffic starts flowing.

Live Analytics & Rule Tuning

Every request logged — cost, latency, model choice. Routing rules adjust at any moment, with no downtime.

Reaches People Where They Work

Native connectors into CRMs, Slack, Teams, and social channels — the model sits behind the channel, not the other way round.

Composable Sub-Workflows

Work breaks into sub-workflows — a summariser, a classifier, a simple trigger — run independently or chained together.

LLM Tooling Wired In

Function calling, retrieval, and structured output are part of the workflow, not bolted on after the fact.

Monitoring Out of the Box

Every execution tracked automatically — errors, latency, throughput — with no separate observability stack to stand up.

Why Route on Complexity, Not Default to the Best

Frontier models are the right tool for a real share of requests and the wrong tool for most of them. Treating every request the same way means paying frontier prices for work a much cheaper model would have handled just as well. Routing on complexity fixes that without asking anyone to change how they use the system.

The orchestration layer — infrastructure, authentication, and the routing workflows themselves — is built on n8n. That's deliberately as far as the implementation detail goes here: the point isn't the tool, it's that request-level routing turns a flat, expensive default into a cost curve that tracks actual complexity.

Facing a Similar Cost Problem?

Happy to talk through routing strategy, self-hosted model options, or where this pattern does and doesn't apply. Best way to reach me is LinkedIn.