Shubham Upadhyay

Senior Backend Architect

Open to Collaborate
Bengaluru, IN
Connect Channels
AI & Context Engineering 6 min read Oct 2024 ArchMCP

Why Dumping 20 Repositories Into an AI Fails (And What We Built Instead)

How moving from passive prompt stuffing to active on-demand MCP tool retrieval turned 50,000-token hallucination traps into 1,500-token precision context packages.

Shubham Upadhyay

Shubham Upadhyay

Senior Backend & Systems Engineer

01.

The "More Context" Trap: Why 200k Tokens Break Down

When LLMs expanded their context windows to 100k and 200k tokens, the first instinct of almost every developer was: "Great! Now I can clone all 15 microservices in our company, drop them into Cursor or Claude Desktop, and the AI will understand our whole architecture."

In theory, it sounded like magic. In practice, it was a disaster.

First, there is the well-documented "lost-in-the-middle" phenomenon. When an LLM receives 150,000 tokens of raw source files, it pays disproportionate attention to the top and bottom of the prompt. Critical interface boundaries, schema definitions, and validation rules in the middle get completely washed away.

Second, the latency and cost penalties are brutal. A 100k-token prompt takes 20 to 45 seconds to generate a response, and burns through daily rate limits in a handful of prompts.

Third, and most dangerous: hallucinated contracts. When flooded with thousands of lines of boilerplate controllers, ORM models, and utility functions, the model starts mixing up fields, inventing endpoints that don't exist, and confusing Service A's database schema with Service B's. More context didn't make the AI smarter—it made it confused.

"An AI coding assistant doesn't need 50,000 lines of implementation code to know how two services talk to each other. It only needs the contract."

— Shubham Upadhyay
02.

The Paradigm Shift: From Passive Stuffing to Active Retrieval

We realized that an experienced senior engineer does not keep 20 repositories memorized line-by-line. Instead, when working on a cross-service task, they ask very specific, high-precision questions: "Which service owns this route?", "What columns are in this table?", "Who calls this API?" That became the design foundation of ArchMCP.

Passive Context Stuffing (The Broken Way)

  • Dumps 20,000 to 100,000 tokens of raw files into the prompt window before every query.
  • Suffers 25 to 45-second latency per response while burning through token budgets.
  • The model hallucinates because API contracts are buried beneath thousands of lines of boilerplate.
  • Fragile: any updated file requires manually re-copying code or re-cloning repositories.
  • Zero awareness of runtime environment variables, Kafka topic subscriptions, or Docker links.

Active MCP Tool Querying (The ArchMCP Way)

  • The AI starts with zero codebase bloat, holding only a lightweight system prompt describing available MCP tools.
  • When the user asks a question, the AI autonomously queries ArchMCP for only what it needs.
  • Responses are surgical context packages (~500 to 1,500 tokens) focused purely on verified contracts.
  • Sub-second retrieval with zero hallucination: the model works with parsed reality, not guesses.
  • Fully compatible with standard MCP clients: Claude Desktop, Cursor, and Google Antigravity.
03.

The Anatomy of Surgical MCP Tools

In src/archmcp/mcp/tools.py, we designed a targeted tool surface where each function answers a single architectural question with zero noise:

HOW THE AI ASSISTANT QUERIES ARCHITECTURE ON DEMAND
1 find_api_owner → Assistant asks: "Who owns POST /api/v1/orders/checkout?"
2 get_service_apis → Retrieves route signatures & parameters, omitting body logic
3 get_database_schema → Returns table columns, data types & keys without raw SQL files
4 get_service_dependencies → Pulls direct callers, Kafka topics & downstream services
5 get_full_context_package → Assembles compact ~1,500-token prompt guidelines & contracts
04.

Inside the 1,500-Token Context Package

When an assistant needs to perform a complex refactor or write a new inter-service feature, it calls get_full_context_package(service_id).

Instead of dumping the entire repository, ArchMCP's ContextService (src/archmcp/services/context_service.py) compiles a high-density, structured dictionary.

It includes the service identity, verified tech stack, public API route table, database schema definitions, and explicit prompt guidelines outlining upstream and downstream compatibility boundaries.

Here is what the actual payload looks like—under 1,500 tokens, perfectly structured for an LLM to reason over without ambiguity:

src/archmcp/services/context_service.py
Context Package Payload
# ContextService.assemble_service_context(service_id)
{
    "service": {
        "id": "order-service",
        "name": "Order Processing Service",
        "owner": "Checkout Team",
        "language": "Go / Gin",
        "apis": [
            {"method": "POST", "path": "/api/v1/orders", "summary": "Create order"},
            {"method": "GET", "path": "/api/v1/orders/{id}", "summary": "Fetch order details"}
        ],
        "database_tables": [
            {"name": "orders", "columns": ["id", "user_id", "total_amount", "status", "created_at"]}
        ],
        "dependencies": {
            "upstream": ["mobile-api-gateway", "web-storefront"],
            "downstream": ["inventory-service", "payment-service", "notification-service"]
        }
    },
    "recommended_prompt_guidelines": "When modifying order-service, ensure backwards compatibility with upstream callers (mobile-api-gateway, web-storefront) and maintain event schema contracts for downstream consumers (inventory-service, payment-service, notification-service)."
}
05.

Token Economics & Speed: Real Numbers

The difference between stuffing entire repositories into a prompt versus querying an on-demand MCP server is not subtle. The numbers speak for themselves:

The Cost of Dumping Repositories

  • Prompt Token Size: 60,000 to 180,000 tokens per message exchange.
  • Latency: 25 to 45 seconds waiting for initial time-to-first-token (TTFT).
  • Cost per 10-turn conversation: $3.00 to $8.00 on frontier models (Claude 3.5 Sonnet / GPT-4o).
  • Rate Limits: Quickly triggers token-per-minute (TPM) throttling on developer accounts.
  • Context Loss: Important edge cases get dropped as prompt history rolls over.

The Economy of Surgical MCP Retrieval

  • Prompt Token Size: ~1,500 to 2,500 tokens per targeted context request.
  • Latency: Sub-second MCP tool execution, 2 to 4 seconds total assistant response time.
  • Cost per 10-turn conversation: Less than $0.15—a 95%+ cost reduction.
  • Zero TPM Throttling: Fits comfortably within standard rate limits.
  • High Accuracy: The model focuses only on relevant contracts, completely eliminating hallucinated routes.
06.

Engineering Takeaways

Precision Trumps Prompt Volume

Giving an AI model 1,500 tokens of verified, structured contracts produces far higher quality code than giving it 100,000 tokens of unindexed source files.

Design Tools Around Human Questions

The best MCP tools mirror the exact questions a developer asks in Slack: "Who owns this endpoint?", "What are the columns here?", "What breaks if I change this?"

Contracts Are All the AI Needs for Integration

When connecting two services, the AI does not need to read the internal business logic of Service B. It only needs the route signature and data contract.

MCP is the Bridge Between Tools and Models

Model Context Protocol allows developer utilities to remain lightweight and headless, turning command-line tools into interactive brains for any AI editor.

Explore Further

Want to review the full ArchMCP overview?

Check out technical architecture, problems faced, and complete tech stack.

View Project Overview
Chat on WhatsApp
Navigating...