Shubham Upadhyay

Senior Backend Architect

Open to Collaborate
Bengaluru, IN
Connect Channels
System Architecture & DX 6 min read Aug 2024 ArchMCP

Why We Skipped Neo4j & Postgres: Building a Zero-Dependency, In-Memory Architecture Index

How a deliberate refusal to over-engineer made ArchMCP fast, zero-dependency, and instantly runnable in under 3 seconds.

Shubham Upadhyay

Shubham Upadhyay

Senior Backend & Systems Engineer

01.

The Default Reflex: "It's a Graph, So You Need Neo4j"

When you set out to build a tool that maps microservices, API routes, database tables, and message queues to give AI coding assistants architecture awareness, the immediate reaction from almost every engineer or architecture forum is unanimous:

"You have microservices communicating with each other. That is a graph problem. You need a dedicated graph database like Neo4j, paired with PostgreSQL for structured metadata, and maybe Elasticsearch or Meilisearch for keyword retrieval."

On paper, that sounds textbook correct. Microservices form a directed network of dependencies. Calling services are nodes, HTTP routes and message queues are edges, and database tables are data sinks. Reaching for Cypher queries, graph traversals, and multi-node clusters seems like the standard, professional thing to do.

The problem starts the exact second you think about developer adoption.

The quickest way to kill a developer tool is to demand that engineers set up infrastructure before they can even try it. If running ArchMCP requires a developer to install Docker, pull a 1.5 GB Neo4j image, configure Java memory heaps, stand up a PostgreSQL container, run schema migrations, and configure connection strings, 95% of developers will close the tab and never look back.

I wanted ArchMCP to feel like ripgrep or jq—you point it at a folder of code, run a command, and it just works in two seconds flat. That single constraint ruled out external database servers completely.

"The fastest way to kill a developer tool's adoption is to make engineers configure three databases before they can run a simple CLI scan."

— Shubham Upadhyay
02.

The Math Behind Architecture Graphs: Why Big Data Fails Here

Before committing to graph databases, we stepped back and actually calculated what we were storing. Big-data graph databases like Neo4j and Amazon Neptune are engineered for social networks with billions of friend edges or fraud-detection clusters analyzing terabytes of transactions. Microservice topologies operate on a completely different scale.

The Heavy Database Approach (Neo4j + Postgres)

  • Neo4j requires an active Java Virtual Machine (JVM) consuming 1.2 GB to 2 GB of idle RAM just to stay online.
  • Demands persistent disk volume mounts, database connection pools, health check retry loops, and schema migration files.
  • Makes running ArchMCP as an ephemeral CLI command (archmcp scan) or local editor subprocess agonizingly slow.
  • Imposes high operational friction for teams deploying remotely—requiring dedicated database instances, backups, and network peering.
  • Network roundtrips to an external database add 5 to 20 milliseconds of latency to every single MCP tool call an AI assistant makes.

The In-Memory Reality (Python Dictionaries & RLock)

  • An entire enterprise architecture of 100 microservices, 3,000 API endpoints, and 500 database tables consumes less than 25 MB of RAM.
  • Direct memory retrieval by service_id or table_name is an O(1) dictionary lookup taking under 5 microseconds.
  • Python's threading.RLock() guarantees thread-safety across concurrent SSE client connections and background queries.
  • Zero installation friction: running pip install archmcp or docker run -p 8000:8000 archmcp boots the entire server in under 2 seconds.
  • Runs natively inside local AI assistants (Claude Desktop, Cursor) via stdio with zero external daemons.
03.

How the In-Memory Store and Fast Token Ranker Actually Work

Rather than offloading queries to external database daemons, ArchMCP orchestrates everything inside the Python process using two specialized components: MetadataDatabase (which holds parsed models and documents) and SimpleSearchStore (which indexes tokens for instant full-text retrieval).

HOW ARCHITECTURE IS DISCOVERED, INDEXED, AND SERVED IN MEMORY
1 Polyglot Discovery → ProjectDetector & RouteExtractor inspect Python, Go, Node & Java repos
2 Thread-Safe Dict Store → MetadataDatabase upserts ServiceMetadata behind threading.RLock()
3 Token Frequency Search → SimpleSearchStore tokenizes endpoints, tables, queues & jobs
4 In-Memory Graph Construction → DependencyLinker maps callers, queues & shared DBs
5 Sub-Millisecond Query Response → MCP Tools return high-precision context packages to AI
04.

Zero-Leak Security: Masking Secrets Before They Reach the AI

One of the most dangerous side effects of building an architecture intelligence tool is credential leakage. When ArchMCP scans a codebase, it naturally encounters .env files, application.properties, and docker-compose.yml containing production API keys and database passwords. If those secrets get indexed or fed into an AI assistant's prompt context, you have created a severe security leak.

Automated Secret Detection & Masking

  • ConfigExtractor actively scans environment files and code references (os.getenv, process.env, os.Getenv).
  • Every discovered key is matched against a security blacklist: secret, password, token, key, auth, credential, jwt, private.
  • Sensitive values are immediately redacted to "********" during the extraction phase before ever being stored in memory or sent to an LLM.
  • Developers can safely ask assistants to inspect configuration flags without leaking live database passwords or third-party API credentials.

Constant-Time Hashed Key Store

  • For remote team deployments, KeyStore manages API keys with tenant scoping and rate limits (60 requests per minute).
  • Uses SHA-256 cryptographic hashing with constant_time_verify() to eliminate timing side-channel attacks.
  • Persists credentials atomically using temporary file writes and os.replace to prevent corrupted state during server crashes.
  • Plaintext keys are displayed only once upon generation and never stored anywhere in memory or on disk.
05.

Graph Traversal in Microseconds: BFS Without a Query Language

When engineers hear that ArchMCP has no graph database, their immediate question is: "How do you calculate transitive blast radiuses and upstream dependencies without Cypher or recursive SQL queries?"

The answer is surprisingly elegant: an adjacency list and Breadth-First Search (BFS) using Python's standard collections.deque.

When a developer or AI assistant asks: "If I touch payment-service or alter the /refund endpoint, what breaks downstream?", ArchMCP builds a reverse caller graph from the in-memory service index in microseconds. It pushes the target service to a queue, walks through direct callers, and recursively discovers every indirect multi-hop dependent across the organization.

Along the way, it collects affected engineering teams, lists impacted endpoints, assigns an impact severity (LOW, MEDIUM, HIGH, or CRITICAL), and generates clean Mermaid sequence diagrams for the assistant to reason over.

src/archmcp/services/architecture_service.py
Blast Radius BFS Engine
# Build reverse caller adjacency graph (callee -> list of callers)
caller_graph: Dict[str, Set[str]] = {s_id: set() for s_id in all_services}
for s_id, s_data in all_services.items():
    for upstream_id in s_data.dependencies.upstream:
        if upstream_id in caller_graph:
            caller_graph[upstream_id].add(s_id)

# BFS to discover direct and transitive multi-hop callers
direct_callers: Set[str] = set(caller_graph.get(service_id, set()))
transitive_callers: Set[str] = set()
visited: Set[str] = {service_id}.union(direct_callers)
queue = deque(list(direct_callers))

while queue:
    curr = queue.popleft()
    for next_caller in caller_graph.get(curr, set()):
        if next_caller not in visited:
            visited.add(next_caller)
            transitive_callers.add(next_caller)
            queue.append(next_caller)
06.

The Payoff: Sub-Second Boots and Frictionless Stdio Mode

Choosing a zero-database, in-memory architecture completely transformed the developer experience of ArchMCP. Instead of being an enterprise platform you have to lobby DevOps to install, it is a tool any developer can run in seconds.

Local Editor Integration (Stdio Mode)

  • Claude Desktop, Cursor, and Antigravity can launch ArchMCP as a native subprocess via stdio with zero background daemons.
  • Scans a monorepo or collection of 10 microservices in 2 to 4 seconds on a standard developer laptop.
  • Consumes approximately 35 MB of total process memory, keeping developer machines cool and responsive.
  • No orphan Docker containers, stale socket files, or zombie database processes left behind when you close your editor.

Remote Team Deployment (ASGI Mode)

  • Deploys as a single lightweight Docker container running Starlette + Uvicorn over Server-Sent Events (SSE).
  • Multi-tenant partitioning (tenant_id) is built directly into the memory store, keeping different teams isolated.
  • Instant horizontal recreation: containers can restart or scale without waiting for database schema migrations.
  • Verified and scored by M8ven for architectural security, low operational footprint, and clean MCP protocol compliance.
07.

Engineering Takeaways

Know Your Data Scale Before Picking Tools

A 25 MB dataset does not need a 1.5 GB graph database cluster. Understanding the real-world scale of your problem prevents massive over-engineering.

Developer Experience (DX) is an Architectural Constraint

If a developer tool requires 15 minutes of database configuration to run locally, developers will abandon it. Zero-setup tools win on adoption every time.

Simplicity Delivers Extreme Speed

Python dictionaries with thread-safe locks and simple BFS queues consistently outperform network calls to external databases by multiple orders of magnitude.

Mask Secrets at the Ingestion Boundary

When building tools for AI coding assistants, security cannot be an afterthought. Redacting secrets before they hit memory or prompts is essential.

Explore Further

Want to review the full ArchMCP overview?

Check out technical architecture, problems faced, and complete tech stack.

View Project Overview
Chat on WhatsApp
Navigating...