Use AI to need less AI.
Coldworks runs your agent workflows — and compiles the judgments they keep repeating into deterministic software you own. Work starts hot and probabilistic. It cools into code. What stays hot is only what genuinely needs a mind.
The industry runs hot on purpose.
Every current product makes agents more numerous, better remembered, more observed, more orchestrated. All of it optimizes the interpreter. Nobody optimizes it away.
There is a reason. Billing by the token makes a vendor's best customer a workflow that never learns.built Coldworks is built on the opposite bet — that agents are a brilliant way to discover a process and an expensive, variable way to run one forever.
Heating — everyone else
- Add more agentsorchestrators orchestrating orchestrators
- Add memoryremember what the agent did, so it can do it again — stochastically
- Add evalsmeasure the misfires without retiring the question
- Grow token spendusage growth is the vendor's KPI, not yours
Cooling — Coldworks
- Record every judgmentdecision, evidence, correction, outcome — append-only
- Compile the repeatsagents write the deterministic replacement
- Guard the boundarycompiled code never runs past its proof
- Route around the modelfalling usage is the product working
Memory products remember what agents did. Eval products measure how agents fail. Coldworks retires the question.
Interpreters
LangGraph, ADK, CrewAI. They execute agent steps — flexibly, expensively, every time. Coldworks runs them inside its hot tier as replaceable plumbing.
Profilers
Langfuse, Braintrust. They watch traces and control nothing. A profiler can tell you a judgment repeated forty times; it cannot route the forty-first.
Dispatch
The layer that decides, per task, whether a model runs at all.spec'd Own the call site and everything else — shadow, audit, attribution — becomes a for-loop.
A JIT compiler for judgment.
Language runtimes solved this thirty years ago: interpret the novel path, profile the hot path, compile it, guard it — and fall back the instant a guard fails.
Coldworks applies the same tiering to decisions. Novel work is interpreted by agents at full flexibility and full price. Repeated work is profiled, compiled into ordinary software, and proven before it takes traffic.built Every compiled path carries the predicate that qualified it, and deoptimizes to the agent the moment reality steps outside it.spec'd
Dispatch — live loop idle
Send tasks through the loop and watch it cool —
including the two moments where it refuses to.
If code can do it, code does it. Agents write the code. Evidence proves it. You own it.
One runtime. One product you can install today.
Coldworks is the parent: the dispatch loop, the guard executor, the crystallizer, the receipt. Doug is its child — the door judgment actually walks through, where judgment happens at volume and every call leaves a trace with its evidence attached. The engine is what turns that stream into code you keep.built
Doug is a warm start, not a cold one. It is our code, making our choices — the answer key, not the exam. Coldworks has to enter systems it did not build, and that is not yet proven.unproven
Cooling without lying.
A machine that is confidently wrong is worse than an agent that is occasionally wrong. Nothing takes traffic on frequency alone: a cluster becomes a candidate only when a discriminator separates its members from nearby-safe neighbours that look like violations and aren't.
01Sealed replay
Candidates are replayed against sealed historical cases — including nearby-safe examples and paired mutations that must flip the result both ways.built
02Shadow before traffic
Qualified artifacts run beside the agent without changing anything, and every disagreement is recorded against the exact candidate version.spec'd
03Unproven is an answer
When evidence can't separate violations from legitimate neighbours, the product says so and keeps that class of work hot. No reassuring green.built
04No self-promotion
Agents propose artifacts and hunt their own counterexamples. They never activate their own output. Promotion is a signed human decision with a rollback path.built
05Permanent audit
A sample of compiled decisions is re-run through the model forever. Drift triggers deopt, agreement is published, and the same machinery is what makes flat pricing safe to offer.spec'd
06Rules can be wrong
Exception growth is read two ways: the repo drifted, or the rule was wrong from inception. Refuting a promoted rule is a first-class outcome, not bypass abuse.spec'd
Agents propose. Evidence qualifies. Maintainers install.
Flat price. Falling temperature.
You pay the same per verified decision on day one and day four hundred. Behind that price, expensive inference gives way to cheap compiled code. The margin that opens is the compilation dividend — we profit only by making your workflow need less AI.unproven
- You own everything compiledPlain code, schemas, tests, fixtures and receipts — readable by your team, in your repo.
- It runs without usExport it, fire us, keep it all. The executor is open and the spec is versioned so third parties can verify the audit claims.spec'd
- Every guard ships with a subtractionPromotion removes the guard's population from the agent's charter — full skips, narrowed prompts, audit-only residue. No subtraction, no promotion.built
- The floor is not zeroNovelty keeps arriving and the audit never stops. T asymptotes to novelty plus audit — we publish the floor instead of pretending it away.
Pay for AI you can safely stop using.
This page has a temperature.
A product that publishes its unproven claims has no business hiding them in its marketing. Every claim above carries a receipt; the ledger below is generated from those receipts, so the page and its evidence cannot drift apart. Click any row to read what backs it.