A runtime for agentic work · working name, clearance pending

Use AI to need less AI.

Coldworks runs your agent workflows — and compiles the judgments they keep repeating into deterministic software you own. Work starts hot and probabilistic. It cools into code. What stays hot is only what genuinely needs a mind.

coldwork (v.) — to strengthen material by working it below the recrystallization temperature: hardening without heat. T is the fraction of decisions resolved by sampling a model. Deterministic code runs at T = 0.
The mark, running6 × 6 · seed 63
molten · sampled coldworked · T = 0
Same units, new arrangement. It crystallizes as you read — the logo is the mechanism. The molten units never settle; the coldworked ones have not moved since they fused. Scroll it, then run your cursor through it.
T ≈ 1.00  ·  The inversion

The industry runs hot on purpose.

Every current product makes agents more numerous, better remembered, more observed, more orchestrated. All of it optimizes the interpreter. Nobody optimizes it away.

There is a reason. Billing by the token makes a vendor's best customer a workflow that never learns.built Coldworks is built on the opposite bet — that agents are a brilliant way to discover a process and an expensive, variable way to run one forever.

Heating — everyone else

  • Add more agentsorchestrators orchestrating orchestrators
  • Add memoryremember what the agent did, so it can do it again — stochastically
  • Add evalsmeasure the misfires without retiring the question
  • Grow token spendusage growth is the vendor's KPI, not yours

Cooling — Coldworks

  • Record every judgmentdecision, evidence, correction, outcome — append-only
  • Compile the repeatsagents write the deterministic replacement
  • Guard the boundarycompiled code never runs past its proof
  • Route around the modelfalling usage is the product working

Memory products remember what agents did. Eval products measure how agents fail. Coldworks retires the question.

Occupied

Interpreters

LangGraph, ADK, CrewAI. They execute agent steps — flexibly, expensively, every time. Coldworks runs them inside its hot tier as replaceable plumbing.

Occupied

Profilers

Langfuse, Braintrust. They watch traces and control nothing. A profiler can tell you a judgment repeated forty times; it cannot route the forty-first.

Open — this is us

Dispatch

The layer that decides, per task, whether a model runs at all.spec'd Own the call site and everything else — shadow, audit, attribution — becomes a for-loop.

T ≈ 0.74  ·  The mechanism

A JIT compiler for judgment.

Language runtimes solved this thirty years ago: interpret the novel path, profile the hot path, compile it, guard it — and fall back the instant a guard fails.

Coldworks applies the same tiering to decisions. Novel work is interpreted by agents at full flexibility and full price. Repeated work is profiled, compiled into ordinary software, and proven before it takes traffic.built Every compiled path carries the predicate that qualified it, and deoptimizes to the agent the moment reality steps outside it.spec'd

Dispatch — live loop idle

Share of decisions by tier, per batchno batches yet
Workflow temperature
1.00
Compiled coverageguards: 0
0%
Cost / decisionflat to you
$0.41
Tier mix0 runs
compiled agent human
Nothing has run yet.
Send tasks through the loop and watch it cool —
including the two moments where it refuses to.
Illustrative simulation, not a live tenant.unproven The behaviour it demonstrates — promotion gated on a human signature, UNPROVEN as a terminal answer, deopt on guard failure — is the actual specified semantics.

If code can do it, code does it. Agents write the code. Evidence proves it. You own it.

T ≈ 0.48  ·  The family

One runtime. One product you can install today.

Coldworks is the parent: the dispatch loop, the guard executor, the crystallizer, the receipt. Doug is its child — the door judgment actually walks through, where judgment happens at volume and every call leaves a trace with its evidence attached. The engine is what turns that stream into code you keep.built

Coldworks — the runtime
Child · Apache-2T ≈ 1.0
Doug shipping
The review agent that generates the judgment stream.

A GitHub App you install on your own keys. It reviews pull requests, executes guards at the CI boundary, and emits a receipt for every call it makes. Free forever, because review bots should be free — you pay for making the bot unnecessary.built

Install Doug →
Parent · commercialT = 0
Engine in build
The crystallizer that turns both records into code you own.

Clusters the repeats, drafts the artifact with its fixtures and discriminator, replays it against sealed history, runs it in shadow — then stops and waits for a maintainer to sign.built

See it run ↑
Doug → engineEvery review decision arrives as a structured trace with its evidence attached.
Engine → DougCompiled guards run inside Doug at the CI boundary, each citing the repeats that qualified it.

Doug is a warm start, not a cold one. It is our code, making our choices — the answer key, not the exam. Coldworks has to enter systems it did not build, and that is not yet proven.unproven

T ≈ 0.22  ·  Trust

Cooling without lying.

A machine that is confidently wrong is worse than an agent that is occasionally wrong. Nothing takes traffic on frequency alone: a cluster becomes a candidate only when a discriminator separates its members from nearby-safe neighbours that look like violations and aren't.

01Sealed replay

Candidates are replayed against sealed historical cases — including nearby-safe examples and paired mutations that must flip the result both ways.built

02Shadow before traffic

Qualified artifacts run beside the agent without changing anything, and every disagreement is recorded against the exact candidate version.spec'd

03Unproven is an answer

When evidence can't separate violations from legitimate neighbours, the product says so and keeps that class of work hot. No reassuring green.built

04No self-promotion

Agents propose artifacts and hunt their own counterexamples. They never activate their own output. Promotion is a signed human decision with a rollback path.built

05Permanent audit

A sample of compiled decisions is re-run through the model forever. Drift triggers deopt, agreement is published, and the same machinery is what makes flat pricing safe to offer.spec'd

06Rules can be wrong

Exception growth is read two ways: the repo drifted, or the rule was wrong from inception. Refuting a promoted rule is a first-class outcome, not bypass abuse.spec'd

Agents propose. Evidence qualifies. Maintainers install.

T ≈ 0.08  ·  Ownership & pricing

Flat price. Falling temperature.

You pay the same per verified decision on day one and day four hundred. Behind that price, expensive inference gives way to cheap compiled code. The margin that opens is the compilation dividend — we profit only by making your workflow need less AI.unproven

Unit economics of one workflow · illustrative
WHAT YOU PAY — FLAT, PER VERIFIED DECISION COMPILATION DIVIDEND WHAT IT COSTS US DAY 1 DAY 400
Token vendors get paid more when you use more. We get paid the same, and win by making the model unnecessary. The obvious failure mode — over-compiling to save tokens — is exactly what the permanent audit stream exists to catch.
  • You own everything compiledPlain code, schemas, tests, fixtures and receipts — readable by your team, in your repo.
  • It runs without usExport it, fire us, keep it all. The executor is open and the spec is versioned so third parties can verify the audit claims.spec'd
  • Every guard ships with a subtractionPromotion removes the guard's population from the agent's charter — full skips, narrowed prompts, audit-only residue. No subtraction, no promotion.built
  • The floor is not zeroNovelty keeps arriving and the audit never stops. T asymptotes to novelty plus audit — we publish the floor instead of pretending it away.

Pay for AI you can safely stop using.

T = 0  ·  Receipts

This page has a temperature.

A product that publishes its unproven claims has no business hiding them in its marketing. Every claim above carries a receipt; the ledger below is generated from those receipts, so the page and its evidence cannot drift apart. Click any row to read what backs it.

0 built — running in the repo today 0 spec'd — schema-frozen, not yet load-proven 0 unproven — no evidence either way
StatusClaimReading