What is harness engineering? It’s about how agent systems wire tools, orchestrate tasks, and manage runtime context.
Here is a repo, a clean-room reverse-engineering project that reconstructs the core patterns behind how Claude Code does harness engineering. GitHub – instructkr/claw-code: The fastest repo in history to surpass 50K stars ⭐, reaching the milestone in just 2 hours after publication. Better Harness Tools, not merely storing the archive of leaked Claude Code but also make real things done. Now rewriting in Rust. · GitHub.
Let Claude Code to summarize how it harness engineering. There are 5 Core Layers of Claude Code’s Harness:
Layer 1 — Registry (What can run?)
Files: src/tools.py, src/commands.py
Tools and commands are loaded once from JSON snapshots into immutable frozen tuples, cached via @lru_cache. Nothing is dynamically registered at runtime.
tools_snapshot.json → load_tool_snapshot() → PORTED_TOOLS (frozen tuple)
commands_snapshot.json → load_command_snapshot() → PORTED_COMMANDS (frozen tuple)
Permission filtering is applied on top — a ToolPermissionContext holds deny-lists (by name or prefix) and is applied at query time:
Layer 2 — Routing (What should run for this prompt?)
File: src/runtime.py → route_prompt()
Token-based scoring — NOT semantic/vector search. The prompt is split into tokens, each token is matched against a tool’s name, source_hint, and responsibility fields.
Top-scoring command + top-scoring tool are selected first, then remaining slots filled.
“review MCP tool” → tokens: {review, mcp, tool}
→ BashTool scores 1, MCPTool scores 2 → MCPTool wins
Layer 3 — Execution Loop (How many turns? When to stop?)
File: src/query_engine.py → QueryEnginePort
This is the heartbeat of the harness. Each submit_message() call:
- Guards on max turns — stops with stop_reason=’max_turns_reached’
- Guards on token budget — stops with stop_reason=’max_budget_reached’
- Records the turn in mutable_messages and transcript_store
- Compacts history after N turns (sliding window — keeps last 10 by default)
- Returns a streaming event sequence: message_start → command_match → tool_match → permission_denial → message_delta → message_stop Config defaults:
max_turns: 8
max_budget_tokens: 2000
compact_after_turns: 12
Layer 4 — State Management (What persists across turns?)
Three separate persistence concerns:

Sessions are fully round-trip serializable — you can save_session() then load_session(session_id) later.
Layer 5 — Bootstrap Graph (How does it all start?)
File: src/runtime.py → bootstrap_session()
The bootstrap executes in a fixed sequence of stages:
prefetch → context snapshot → setup/trusted init
→ execution registry build
→ route_prompt()
→ execute matched commands/tools
→ infer permission denials
→ stream_submit_message()
→ persist_session()
✶ Insight
The bootstrap graph is why Claude Code feels “instant” — context, registry, and routing all happen before the first model call. The harness front-loads everything deterministic
so the LLM only handles what requires intelligence.
Full Data Flow (one turn)
User prompt
│
▼
route_prompt() ← token-score against tool/command registry
│
▼
filter_tools_by_permission_context() ← deny lists applied
│
▼
build_execution_registry() ← dispatch layer
│
▼
QueryEnginePort.submit_message() ← guards, state mutation, usage tracking
│
├── stream events (UI layer)
└── TurnResult (stop_reason, usage, matched tools/commands)
│
▼
session_store.save_session() ← persist JSON