Deploying MedGamma Mini Model Locally and Create an MCP Server

I have an idea to let MedGemma become the specialized local engine, MCP becomes the interface, and Claude/ChatGPT becomes the high-level reasoning/orchestration layer. The flow would be Claude Desktop sends a tool call → the MCP server builds a prompt → llama-server runs MedGemma inference on CPU → structured JSON comes back to Claude.

Deploying locally can effectively maintain privacy and reduce costs; however, a notable disadvantage is that CPU inference tends to be slower than that of a GPU. Nonetheless, for structured extraction tasks that do not require real-time interaction, a rate of 10 tokens per second is entirely adequate.

Several key decisions to make/made:

1. llama.cpp + GGUF, not transformers

Running transformers with a 4B model on CPU gives you 1-3 tokens/second and multi-minute generations. That’s unusable.

llama.cpp is purpose-built for CPU inference. It uses GGUF format (quantized model weights) and includes architecture-specific CPU backends — on our Intel i7-12800H, it auto-selected the ggml-cpu-alderlake.dll backend for Alder Lake optimizations.

Result: ~11 tokens/second for a 4B Q4_K_M model on a 14-core CPU. A 200-token JSON extraction takes ~18 seconds. Not fast, but practical.

2. Q4_K_M quantization

The model is 4B parameters. In full precision (bf16), that’s ~8GB. On CPU, you want it smaller — both for RAM and for inference speed (less data to move).

QuantSizeRAMAccuracy RetainedVerdict
Q4_K_M2.5GB~3GB~81%Sweet spot for CPU
Q8_04.1GB~5GB~99%If you have RAM to spare
IQ2_XS0.85GB~1.9GB~63%Only if RAM-starved

We went with Q4_K_M (2.5GB). On a 32GB RAM machine, Q8_0 would have been fine too, but Q4_K_M loads faster and the accuracy loss is minimal for structured extraction tasks where the model is following a clear JSON schema.

3. Two-process architecture

We did not load the model inside the MCP server process. Instead:

  • llama-server (from llama.cpp) loads the GGUF and exposes an OpenAI-compatible HTTP endpoint (/v1/chat/completions) on localhost
  • FastMCP server is a thin Python HTTP client — each tool builds a prompt, calls llama-server, parses the JSON response

Why separate? If the model crashes or needs restarting, the MCP server stays up. If you want to swap models (e.g. try Q8_0), you restart llama-server only. The MCP server doesn’t care what’s behind the HTTP endpoint.

4. FastMCP with stdio transport

The MCP server uses FastMCP (same framework as my production FactSet index plugin), but with stdio transport instead of HTTP. stdio is the standard for local MCP servers consumed by Claude Desktop — Claude launches the server process, communicates over stdin/stdout, and manages its lifecycle.

5. Text-only first, vision later

MedGemma 1.5 4B is multimodal (text + medical images). We skipped the vision projector (mmproj) for the initial build because:

  • EHR extraction is a text task
  • llama.cpp’s Gemma 3 vision support has known issues on Windows
  • The mmproj file adds ~1GB to RAM usage

Vision is addable later — download the mmproj file, add --mmproj to the llama-server command, and add an image-encoding tool.

How is it done? No need to build from source. llama.cpp publishes prebuilt Windows binaries on their GitHub releases. We downloaded llama-b10809-bin-win-cpu-x64.zip (17.6MB) and extracted to medgemma-mcp/bin/. The zip includes llama-server.exe plus architecture-specific DLLs — llama.cpp auto-selects the best one at runtime.

MedGemma is a gated model, so I do need to login huggingface, create an account, create access token(read-only), login and download the unsloth/medgemma-1.5-4b-it-GGUF quantized version — same model, pre-converted to GGUF.

$env:HF_TOKEN = "hf_token"
python -m huggingface_hub.commands.huggingface_cli login --token $env:HF_TOKEN
.\scripts\download_model.ps1 # pulls Q4_K_M, ~2.5GB

Then Start llama-server

.\bin\llama-server.exe `
-m .\models\medgemma-1.5-4b-it-Q4_K_M.gguf `
--host 127.0.0.1 --port 8080 `
--alias medgemma-1.5-4b `
-c 8192 -t 14 -ngl 0
curl http://127.0.0.1:8080/health
# {"status":"ok"}

Configure Claude Desktop:

"mcpServers": {
"medgemma": {
"command": "python",
"args": [
"C:\\Users\\you\\medgemma-mcp\\server\\app.py"
]
}
}

Five tools are built:

extract_lab_values
Input: lab report / EHR text → Output: JSON array (test_name, value, unit, reference_range, abnormal_flag)
extract_medications
Input: med list / discharge summary → Output: JSON array (medication, dose, route, frequency, status)
extract_diagnoses
Input: problem list / encounter note → Output: JSON array (diagnosis, icd_code, status)
summarize_ehr
Input: encounter note → Output: 3-5 sentence summary
extract_structured
Input: document text + schema description → Output: JSON matching the described schema

Resources

Leave a Reply