Harness Engineering

What is harness engineering? It’s about how agent systems wire tools, orchestrate tasks, and manage runtime context.

Here is a repo, a clean-room reverse-engineering project that reconstructs the core patterns behind how Claude Code does harness engineering. GitHub – instructkr/claw-code: The fastest repo in history to surpass 50K stars ⭐, reaching the milestone in just 2 hours after publication. Better Harness Tools, not merely storing the archive of leaked Claude Code but also make real things done. Now rewriting in Rust. · GitHub.

Let Claude Code to summarize how it harness engineering. There are 5 Core Layers of Claude Code’s Harness:


Layer 1 — Registry (What can run?)

Files: src/tools.py, src/commands.py

Tools and commands are loaded once from JSON snapshots into immutable frozen tuples, cached via @lru_cache. Nothing is dynamically registered at runtime.

tools_snapshot.json → load_tool_snapshot() → PORTED_TOOLS (frozen tuple)
commands_snapshot.json → load_command_snapshot() → PORTED_COMMANDS (frozen tuple)

Permission filtering is applied on top — a ToolPermissionContext holds deny-lists (by name or prefix) and is applied at query time:


Layer 2 — Routing (What should run for this prompt?)

File: src/runtime.py → route_prompt()

Token-based scoring — NOT semantic/vector search. The prompt is split into tokens, each token is matched against a tool’s name, source_hint, and responsibility fields.
Top-scoring command + top-scoring tool are selected first, then remaining slots filled.

“review MCP tool” → tokens: {review, mcp, tool}
→ BashTool scores 1, MCPTool scores 2 → MCPTool wins


Layer 3 — Execution Loop (How many turns? When to stop?)

File: src/query_engine.py → QueryEnginePort

This is the heartbeat of the harness. Each submit_message() call:

  1. Guards on max turns — stops with stop_reason=’max_turns_reached’
  2. Guards on token budget — stops with stop_reason=’max_budget_reached’
  3. Records the turn in mutable_messages and transcript_store
  4. Compacts history after N turns (sliding window — keeps last 10 by default)
  5. Returns a streaming event sequence: message_start → command_match → tool_match → permission_denial → message_delta → message_stop Config defaults:
    max_turns: 8
    max_budget_tokens: 2000
    compact_after_turns: 12

Layer 4 — State Management (What persists across turns?)

Three separate persistence concerns:

Sessions are fully round-trip serializable — you can save_session() then load_session(session_id) later.


Layer 5 — Bootstrap Graph (How does it all start?)

File: src/runtime.py → bootstrap_session()

The bootstrap executes in a fixed sequence of stages:

prefetch → context snapshot → setup/trusted init
→ execution registry build
→ route_prompt()
→ execute matched commands/tools
→ infer permission denials
→ stream_submit_message()
→ persist_session()

✶ Insight
The bootstrap graph is why Claude Code feels “instant” — context, registry, and routing all happen before the first model call. The harness front-loads everything deterministic
so the LLM only handles what requires intelligence.


Full Data Flow (one turn)

User prompt
│
▼
route_prompt() ← token-score against tool/command registry
│
▼
filter_tools_by_permission_context() ← deny lists applied
│
▼
build_execution_registry() ← dispatch layer
│
▼
QueryEnginePort.submit_message() ← guards, state mutation, usage tracking
│
├── stream events (UI layer)
└── TurnResult (stop_reason, usage, matched tools/commands)
│
▼
session_store.save_session() ← persist JSON

Leave a Reply