Stop chaining tool calls.
Let the model write code.
Code Mode converts MCP tools into a typed TypeScript API and lets the model write a single program that calls them. One inference, many tool calls, no round-trips through the model's context. Five to ten times cheaper. Often faster too.
The old way is wasteful.
Most agents expose MCP tools to the LLM as function calls. The model picks one, you run it, you stuff the result back into the prompt, the model picks the next one. Every step is a full round-trip, every intermediate result is paid for in tokens.
// Traditional MCP, 5 tool calls = 5 inferences const a = await model.tool_call("weather", { city: "Austin" }); // 1500 tokens written back to context const b = await model.tool_call("weather", { city: "Reykjavik" }); // 1500 more tokens const answer = a.tempC < b.tempC ? a : b; // answer back to context
Code Mode flips it. The model writes a TypeScript program once, the program does all the work, only the final answer goes back to the model. Cloudflare benchmarked it at 70-90% fewer tokens for multi-step tasks.
// Code Mode. 1 inference, N tool calls const [a, b] = await Promise.all([ mcp.weather.get({ city: "Austin" }), mcp.weather.get({ city: "Reykjavik" }), ]); console.log(a.tempC < b.tempC ? a : b);
Why it works.
LLMs have seen real code.
Trained on millions of TypeScript repos. Tool-call JSON is a synthetic format they only saw in fine-tuning. Writing code is their native mode.
Loops, conditionals, batching.
Want to call the same MCP tool 50 times in parallel? Trivial in code. A nightmare in raw tool calling.
Only the answer comes back.
Intermediate values stay in the sandbox. The model never reads the JSON it doesn't need. Tokens saved, latency saved, money saved.
Same task. Both ways.
5 weather lookups + ranking
5 weather lookups + ranking
Synthetic benchmark using OpenAI gpt-4o-mini calling a mock weather MCP. Your mileage will vary by task; the larger the fan-out, the larger the win.
Five lines of SDK.
import { OmniStrat } from "@omnistrat/sdk"; const client = new OmniStrat({ apiKey: process.env.OMNISTRAT_KEY!, orgId: "my-org" }); const r = await client.codeMode.execute({ prompt: "Get current weather in Austin and Reykjavik. Return whichever is colder.", mcp_servers: [{ name: "weather", url: "https://weather.mcp.example/" }], }); console.log(r.result); // { city: "Reykjavik", tempC: 4 } console.log(r.code); // the TS program the model wrote
Safety is built in.
The generated code runs in a sandboxed isolate with no network access. The only things it can reach are the MCP servers you pass in, and even those are reached through a proxy this Router holds the credentials for. The model cannot write code that leaks an API key, because the API key is never in the sandbox.
Cloudflare calls this bindings, not network. We call it the right way.
Built on the same primitive as Cloudflare's.
It turns out we've all been using MCP wrong. LLMs are better at writing code to call MCP than at calling MCP directly. , Kenton Varda & Sunil Pai, Cloudflare, September 2025
We built OmniStrat Code Mode on the same isolate primitive Cloudflare pioneered: production-grade sandboxing with the same isolation guarantees, engineered directly into our Router. You don't have to be a Cloudflare customer to use it; OmniStrat handles the entire runtime.
Try Code Mode tonight.
Install the SDK and you have it. Code Mode authenticates with the same orgId and apiKey as every other call, and counts against the same usage and spend caps. Including the hard cap, which refuses the request before a provider is called rather than reporting the overspend afterwards.