Open-source coding agent · your terminal

Your models.
One coding team.

Plan, build and review with the models you choose. Marshall brings local and hosted models into one workflow — with you in control of what runs.

Read the docs ↗

Local-first. Model-flexible. Approval-gated. Install with npm instead. · Browser extension download

From a task to a reviewed change.Illustrative session · not live output
marshall · ~/projects/storefront
❯ Add validation to checkout and cover it with tests.
◆ planner Map the form, validation rules and existing tests.
model DeepSeek · hosted
● coder Add validation and test the edge cases.
model Qwen · local / llama.cpp
edit_file src/checkout.ts · tests/checkout.test.ts
run_shell npm test
✓ reviewer Check the diff against the original request.
✓ Tests passed. Ready for your review.

measured

Same work. Less overhead.

We ran the same models on the same tasks and graded the result with code, not an LLM judge. Small samples, published with the task and the caveats — but the direction is promising.

63%

fewer input tokens than Codex

Marshall and Codex both passed 2/2 Terminal-Bench tasks with the same GPT-5.6 Luna model and account. Marshall reported 119k input tokens to Codex's 319k.

62%

fewer tool calls than pi

On a 28-file migration with 122 verifier tests, Marshall averaged 5.7 calls to pi's 14.8 on GPT-5.6 Luna. Every reported trial passed: 3/3 and 4/4.

~90%

cheaper cached OpenRouter turns

In a three-turn live API check, prompt caching cut follow-up turns from $0.00955 to $0.00096 and $0.00098. Same 3,746-token prefix, now reused.

These are task-specific results, not a claim that Marshall always wins. See configurations, methodology and limitations.

the team

One job. Several models.

A planner, a coder and a reviewer come for free. Then you add whoever the work needs — a tester that can only read and run, a summariser on a cheap local model, a reviewer on something stubborn. Each one has a name, a model, and a one-line brief.

01

planner

Turns the ask into an ordered plan. Pin it to whatever thinks best.

02

coder

Reads, edits, runs tests. Delegates to named agents when the work splits.

03

reviewer

Looks at the actual diff. Does not rubber-stamp.

Mix however you want. DeepSeek V4 Flash on planning, Qwen3.6 27B writing the code, Kimi K3 reviewing. All local, all hosted, or split across both. /team add walks you through it.

open weights

Built for models you actually run

Most agents treat local inference as a fallback. Marshall starts there. llama.cpp and Ollama are first-class — no key, no account, no cloud.

It sees what's loaded

Which models are resident, real context length, quant, file size. Loaded ones sort first so you don't eat a cold reload mid-task. On OpenRouter, cost per million tokens is right there.

Fast and deep, any vendor

Planning and review on a frontier model. File reads and history compression on a local one. Or both local, on two machines. The bill follows the split.

Match the belt to the model

light for a small model that shouldn't spawn helpers. default for everyday work. agentic when you want named agents running in the background.

Bricks: Floating City — a LEGO-style world running in the browser

Built entirely on a local 27B

A LEGO-style floating city with a zombie-siege mode, every line written by Qwen 3.8 27B through Marshall — no cloud model involved. Play it.

Pagoda in the Garden of Falling Blossoms — a voxel scene running in the browser

Same model, a different mood

A five-tier pagoda on a floating island, 81,000 voxel cubes, day-to-night lighting and a working camera rig — also Qwen 3.8 27B through Marshall, no cloud model involved. Explore it.

safety

Three levels. You pick.

Default is you. Every write and every shell command shows a diff before it runs. Turn the dial when you want something else.

default

you approve

Human in the loop. Approve once, always-allow for the session, or deny.

yolo

hands off

No prompts. For a sandbox or CI — not for a repo you care about.

agentic

a judge

A fast model reviews each tool call. Safe skips you. Unsafe still asks, with the reason attached.

also

The rest, without the lecture

Interrupt and steer

esc stops at the next safe boundary. Your next message course-corrects — no restart, no half-written files.

Project memory

AGENTS.md is loaded at session start. Conventions, architecture, decisions — so a smaller model already knows the house rules.

Background work

Long jobs keep running. History compresses when it has to. The terminal is one client of a headless engine.

Open, all the way down

MIT licensed. Built on agention. No hosted service in the middle.

Leaves nothing behind

--private — no session log, no scratch notes, nothing written to config.json. OpenRouter requests are routed to deny data collection; local models need no such flag.

providers

Your weights, your hardware, your call

Setup leads with llama.cpp, Ollama and OpenRouter. Pair any two. Switch with /model.

llama.cpp · local · no key · probes what's loaded ollama · local · no key · resident models first openai compatible · vLLM, LM Studio, … · --host openrouter · live catalogue · free tiers ranked first claude · ANTHROPIC_API_KEY or login openai · OPENAI_API_KEY gemini · GEMINI_API_KEY mistral · MISTRAL_API_KEY

No GPU? OpenRouter lists models that can actually call tools, with :free variants first. Cost per million tokens sits next to the name.

commands

Small surface, muscle memory

The ones you'll actually type. /help prints the rest.

/team addname an agent, pin a model, give it a one-line brief
/modelpick deep and fast — local what's-loaded, or OpenRouter with cost
/runtimelight · default · agentic — match the tool belt to the model
/safetydefault · yolo · agentic — you, nobody, or a judge
/planget a plan before touching code
escstop and steer. twice to force-quit

fine print

What it doesn't do

Worth knowing before you point it at something that matters.

The sandbox is containment, not a security boundary. File tools stay in the workspace. A shell command can still reach the network or an absolute path.
The model still has to call tools well. Marshall gives every provider the same belt. Whether it uses it is on the weights.
Benchmarks are snapshots, not universal rankings. Results depend on the task, model, harness configuration and provider. We publish the trial counts and the losses too.
One workspace, one terminal. No editor plugin yet. The engine is headless — a VS Code client is on the list.

get started

Your models are already running

If llama.cpp or Ollama is up, marshall finds it. No account. Or install from npm.

star on github