Give AI agents full understanding of your codebase β not just fragments.
CortexHarness extracts real structure from your source code and documentation, stores it in a hybrid Graph + Vector engine, and serves it to AI agents through MCP. The result: agents that reason over your actual architecture instead of guessing from incomplete context.
AI coding agents hallucinate when context is missing. They see a file but not the call graph. They read a function but not the framework convention. They answer questions but can't trace impact.
CortexHarness fixes this by building a persistent, queryable knowledge layer that gives agents:
- Structural truth β call graphs, dependencies, type hierarchies, framework routes (not just text)
- Semantic understanding β embeddings for similarity search across code and docs
- Complete context β one query returns the function, its callers, its framework role, and related documentation
- Less token waste β agents ask the graph instead of stuffing entire files into prompts
Your Codebase (20+ languages) Your Documentation (PDF, DOCX, MD, PPTX)
β β
βΌ βΌ
ββββββββββββββββββββββ ββββββββββββββββββββββ
β Language-Specific β β GraphRAG Pipeline β
β Analyzers β β Entity Extraction β
β Extract: calls, β β Extract: entities,β
β types, flows, β β relationships, β
β framework roles β β summaries β
ββββββββββ¬ββββββββββββ ββββββββββ¬ββββββββββββ
β β
βΌ βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Hybrid Storage (Embedded or Remote) β
β β
β Graph Providers Qdrant β
β βββββββββββββββ βββββββββββββββββββββ β
β β’ FalkorDB (default) β’ Code embeddings β
β β’ LadybugDB (Kuzu-based) β’ Doc embeddings β
β β’ Neo4j (optional) β’ Entity-linked paragraphsβ
β β’ Call graphs, dependencies,β’ Semantic similarity β
β type hierarchies, flows β’ RAG-ready retrieval β
β β
ββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Unified MCP Server β 30+ Query Tools β
β β
β "What calls this function?" β Graph traversal β
β "Find similar implementations" β Vector search β
β "Trace this API end-to-end" β Fullstack chain β
β "What breaks if I change this?" β Impact analysis β
β "Summarize this module" β Graph + vectors β
ββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββ
β
βββββββββββββββΌββββββββββββββ
βΌ βΌ βΌ
Claude Code Qwen Code Cursor / Copilot
Embedded FalkorDB + Qdrant = no Docker, no daemon, no infrastructure. Clone, make install, dev sync code β done. Switch to remote (Docker or cloud) when your team needs shared storage.
make install # one command: venv + deps + global CLI
dev init # interactive wizard
dev sync code # graph + embeddings in one pass
dev start # MCP servers ready for your AI agent- Graph-first pipeline: topology and relationships are written before embeddings. Agents can query structure immediately.
- Incremental sync: Git-aware change detection β SHA-256 content verification β only reprocess what changed. A 500K-line repo syncs in seconds after a one-file commit.
- Framework-aware: Spring controllers, ASP.NET routes, Struts actions, MyBatis mappers β extracted as first-class graph nodes, not buried in raw text.
- Crash-safe ingestion: write-ahead journal for graph mutations. If sync fails mid-write, recovery replays exactly what was interrupted.
AI agents have finite context windows. CortexHarness helps them use it wisely:
| Without CortexHarness | With CortexHarness |
|---|---|
| Agent reads 10 files to find callers | query_subgraph returns the call tree in one tool call |
| Agent guesses the framework pattern | get_framework_context returns the exact overlay |
| Agent stuffs entire modules into prompt | semantic_search returns the 5 relevant paragraphs |
| Agent hallucinates the dependency chain | trace_flow returns the verified path |
Not just ingestion β a complete pipeline from source to agent query:
- Ingest β
dev sync code/dev sync docextracts structure and semantics; doc sync also generates dynamic YAKE entity rules per document before GLiNER extraction - Store β graph relationships + vector embeddings persisted to disk
- Serve β MCP servers expose 30+ query tools on localhost
- Query β AI agents call MCP tools instead of reading raw files
- Update β next sync only processes changes, graph stays current
The ADLC skill pack (make install-adlc) installs 12+ battle-tested skills into your AI agent:
- Plan before code β architecture-first planning, not dive-and-pray
- Debug systematically β root cause analysis with evidence, not guess-and-check
- Predict risks β 5-persona panel reviews changes before implementation
- Generate edge cases β 12-dimension scenario coverage
- Audit security β STRIDE + OWASP with automated fix suggestions
Works with: Claude Code, Qwen Code, OpenCode, GitHub Copilot, Cursor, Continue, Antigravity.
- Python 3.12+
- uv (fast Python package manager)
- Node.js β₯ 18 (only for ADLC skill pack)
- (optional) .NET SDK 8+ and JDK 17+ β needed only for full-fidelity C#/VB.NET (Roslyn) and VB6 (ANTLR) parsing. See Parser Environment Requirements.
git clone https://github.com/baka3k/cortex-harness.git
cd cortex-harness
make install # create .venv, install deps, register global `dev` CLI
make storage-init # initialize persistent storage at ~/.cortext-harness/v1/
make doctor # verify Qdrant + FalkorDB connectivity
dev init # interactive wizard β configure project, backend, folders
dev sync code # ingest source code β graph + embeddings
dev sync doc # ingest documentation β knowledge graph + vectors
dev start # MCP servers: code (:8788) + doc (:8789)Point your AI agent at http://localhost:8788 (code) or http://localhost:8789 (doc). The agent gets full graph + vector query access through standard MCP protocol.
Compiler-grade parsers β not regex, not naive AST. Each analyzer uses the right tool for the language:
| Language | Parser Technology | What It Extracts |
|---|---|---|
| C# / VB.NET | Roslyn workspace (semantic model) | Full type resolution, method invocations, LINQ, async/await, ASP.NET overlays |
| VB6 / VBA | ANTLR4 (whole-program parse) | Call graphs, control flow, designer forms, control arrays |
| Java / Kotlin | Tree-sitter + Spring/Struts/MyBatis overlays | Calls, generics, inheritance, framework routes, DI wiring |
| C / C++ | Clang-based | Pointers, templates, macros, header resolution |
| Go, Rust, Swift | Tree-sitter | Calls, types, interfaces, concurrency patterns |
| TypeScript / JavaScript | Tree-sitter | Imports, exports, async, React/Vue components |
| Python | AST + import graph | Decorators, generators, type hints |
| COBOL | Custom parser (copybook-aware) | Paragraphs, COPY resolution, file I/O |
| Perl, Delphi, Dart | Tree-sitter / custom | Module structure, calls, inheritance |
| Android | Mixed Java + Kotlin analyzer | Cross-language calls, Activity lifecycle |
Framework overlays: Spring, Struts, MyBatis, Flutter, ASP.NET Core, ASP.NET Framework, Express.js, FastAPI/Django, Laravel
Document formats: PDF, Markdown, DOCX, PPTX, XLSX, TXT
Some analyzers depend on an external toolchain. Prepare the ones below to get full-fidelity parsing:
| Environment | Required By | Behavior When Missing |
|---|---|---|
.NET SDK 8+ (dotnet on PATH) |
C# / VB.NET parsers β Roslyn semantic engine | Falls back to tree-sitter (C#) / regex (VB.NET) parsing with a logged warning |
JDK 17+ (java on PATH) |
VB6/VBA parser β ANTLR engine (worker runs as a java -jar process) |
Engine auto falls back to the regex engine with a one-line warning; forcing engine=antlr fails loudly |
| Maven 3.9+ | Rebuilding the VB6 ANTLR worker jar only | Optional at parse time β a prebuilt jar ships with the repo |
| ANTLR4 | VB6/VBA whole-program parse | Nothing to install separately β the grammar is bundled inside the vendored worker jar and runs on the JDK above |
Notes:
- Roslyn workers self-build: the first C#/VB.NET run executes
dotnet buildagainst the bundled worker projects (code-tiny/tools/csharp/roslyn_worker/,code-tiny/tools/vb/roslyn_worker/); a matching prebuilt DLL is reused when available. - Java & Kotlin need no external toolchain: Java sources are parsed with tree-sitter grammars installed via pip (
tree-sitter-java,tree-sitter-kotlin). Parsing Java code does not require a JDK β the JDK above is consumed only by the VB6 ANTLR engine.
Default run vs. standalone runs. The default sync (dev sync code / dev sync code all) scans every configured language analyzer, so a missing toolchain is reported as a warning or per-language error for exactly the parts that need it. When you run an individual analyzer standalone, you only need the toolchain for the languages that analyzer covers β e.g. syncing a Java-only repo requires none of the environments above.
Embedded by default β data lives on disk, no network required.
~/.cortext-harness/v1/instances/default/
βββ manifest.json
βββ qdrant/{code,doc}/ # vector embeddings
βββ falkordb/{code,doc}/data.rdb # graph relationships
βββ backups/
Remote mode for shared team infrastructure:
make infra-up # start Docker: Qdrant (:6333) + FalkorDB (:6379)
dev init # select "remote" backend, enter connection detailsMultiple projects share one storage instance with isolated graph/collection namespaces. Override paths with CORTEX_DATA_HOME or CORTEX_STORAGE_INSTANCE.
| Command | Description |
|---|---|
dev build |
Create/reuse .venv, install dependencies |
dev doctor |
Connectivity probes + MCP port diagnostics |
dev start |
Launch MCP servers (code + doc) |
dev start --server code --name shop --port 8790 |
Named instance with custom port |
dev stop |
Stop all MCP servers |
dev stop --name shop |
Stop one named instance |
| Command | Description |
|---|---|
dev sync code |
Incremental code sync (graph + embeddings) |
dev sync code all |
Full sync across all configured source roots |
dev sync code --sync-mode graph --full-scan |
Graph-only rebuild |
dev sync code --sync-mode embedding --full-scan |
Embedding-only rebuild |
dev sync doc |
Incremental doc sync (GraphRAG + vectors) |
dev sync doc all |
Full doc sync across all configured doc roots |
| Command | Description |
|---|---|
dev init |
Interactive setup wizard |
dev status |
Show active config |
dev ignore add <FOLDER>... |
Add folders to scan-ignore list |
dev install-adlc |
Bootstrap ADLC skill pack for AI agents |
make export-db OUTPUT=project.cortexdb |
Export project database |
make import-db ARCHIVE=project.cortexdb |
Import project database |
One server, every query pattern:
| Category | Key Tools |
|---|---|
| Symbol & graph | search_functions, get_symbol, query_subgraph, explore_graph |
| Flow tracing | trace_flow, reconstruct_flow, find_workflows_containing |
| Impact analysis | analyze_workflow_impact, find_paths, compute_scc |
| Fullstack bridge | get_api_call_chain, find_callers_of_endpoint |
| Semantic search | semantic_search, search_by_code |
| Project context | get_project_modules, get_framework_context, get_public_apis |
| Dependency planning | topological_sort, plan_dependency_order |
cortex-harness/
βββ cortex_harness/ # CLI (dev.py), config, storage management
βββ code-tiny/ # Language analyzers, MCP servers, graph writers
β βββ tools/ # Per-language analyzers (csharp, java, go, rust, ...)
β βββ mcp/ # Unified MCP server + per-language backends
β βββ tools/graph/ # Shared graph schema, writers, CLI helpers
βββ doc-tiny/ # Document pipeline, GraphRAG, entity extraction
βββ scripts/ # Lifecycle, MCP management, migrations
Design principles:
- Embedded-first β zero infrastructure for individual developers
- Provider-neutral β swap FalkorDB β Neo4j without touching analyzers
- Project-scoped β multiple projects, one storage, isolated namespaces
- Incremental by default β Git-aware, SHA-256 verified, only changed files
- Crash-safe β write-ahead journal for graph mutations with automatic replay
Windows: ModuleNotFoundError: No module named 'requests'
uv pip install --python .venv\Scripts\python.exe --requirements code-tiny\requirements.txtWindows: 'dev' is not recognized as a command
Run directly: .venv\Scripts\dev.exe <command>, or add %USERPROFILE%\.local\bin to PATH.
CUDA / PyTorch issues
uv pip install torch torchvision torchaudio --default-index https://download.pytorch.org/whl/cu128
python -c "import torch; print(torch.cuda.is_available())"Port already in use
dev start --port 8790 to pick a different port, or dev stop to clear existing servers.
See LICENSE for details.