Skip to content

Latest commit

Β 

History

417 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

CortexHarness

Give AI agents full understanding of your codebase β€” not just fragments.

CortexHarness extracts real structure from your source code and documentation, stores it in a hybrid Graph + Vector engine, and serves it to AI agents through MCP. The result: agents that reason over your actual architecture instead of guessing from incomplete context.


The Problem It Solves

AI coding agents hallucinate when context is missing. They see a file but not the call graph. They read a function but not the framework convention. They answer questions but can't trace impact.

CortexHarness fixes this by building a persistent, queryable knowledge layer that gives agents:

  • Structural truth β€” call graphs, dependencies, type hierarchies, framework routes (not just text)
  • Semantic understanding β€” embeddings for similarity search across code and docs
  • Complete context β€” one query returns the function, its callers, its framework role, and related documentation
  • Less token waste β€” agents ask the graph instead of stuffing entire files into prompts

How It Works

 Your Codebase (20+ languages)        Your Documentation (PDF, DOCX, MD, PPTX)
          β”‚                                      β”‚
          β–Ό                                      β–Ό
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”               β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚  Language-Specific β”‚               β”‚  GraphRAG Pipeline β”‚
 β”‚  Analyzers         β”‚               β”‚  Entity Extraction β”‚
 β”‚  Extract: calls,   β”‚               β”‚  Extract: entities,β”‚
 β”‚  types, flows,     β”‚               β”‚  relationships,    β”‚
 β”‚  framework roles   β”‚               β”‚  summaries         β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜               β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
          β”‚                                      β”‚
          β–Ό                                      β–Ό
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚           Hybrid Storage (Embedded or Remote)           β”‚
 β”‚                                                         β”‚
 β”‚   Graph Providers             Qdrant                    β”‚
 β”‚   ───────────────             ─────────────────────     β”‚
 β”‚   β€’ FalkorDB (default)        β€’ Code embeddings         β”‚
 β”‚   β€’ LadybugDB (Kuzu-based)    β€’ Doc embeddings          β”‚
 β”‚   β€’ Neo4j (optional)          β€’ Entity-linked paragraphsβ”‚
 β”‚   β€’ Call graphs, dependencies,β€’ Semantic similarity     β”‚
 β”‚     type hierarchies, flows   β€’ RAG-ready retrieval     β”‚
 β”‚                                                         β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          β”‚
                          β–Ό
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚         Unified MCP Server β€” 30+ Query Tools            β”‚
 β”‚                                                         β”‚
 β”‚   "What calls this function?"        β†’ Graph traversal  β”‚
 β”‚   "Find similar implementations"     β†’ Vector search    β”‚
 β”‚   "Trace this API end-to-end"        β†’ Fullstack chain  β”‚
 β”‚   "What breaks if I change this?"    β†’ Impact analysis  β”‚
 β”‚   "Summarize this module"            β†’ Graph + vectors  β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          β”‚
            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
            β–Ό             β–Ό             β–Ό
      Claude Code    Qwen Code    Cursor / Copilot

Why Developers Choose CortexHarness

Zero-Config Start

Embedded FalkorDB + Qdrant = no Docker, no daemon, no infrastructure. Clone, make install, dev sync code β€” done. Switch to remote (Docker or cloud) when your team needs shared storage.

make install          # one command: venv + deps + global CLI
dev init              # interactive wizard
dev sync code         # graph + embeddings in one pass
dev start             # MCP servers ready for your AI agent

Accurate, Not Just Fast

  • Graph-first pipeline: topology and relationships are written before embeddings. Agents can query structure immediately.
  • Incremental sync: Git-aware change detection β†’ SHA-256 content verification β†’ only reprocess what changed. A 500K-line repo syncs in seconds after a one-file commit.
  • Framework-aware: Spring controllers, ASP.NET routes, Struts actions, MyBatis mappers β€” extracted as first-class graph nodes, not buried in raw text.
  • Crash-safe ingestion: write-ahead journal for graph mutations. If sync fails mid-write, recovery replays exactly what was interrupted.

Less Context, More Signal

AI agents have finite context windows. CortexHarness helps them use it wisely:

Without CortexHarness With CortexHarness
Agent reads 10 files to find callers query_subgraph returns the call tree in one tool call
Agent guesses the framework pattern get_framework_context returns the exact overlay
Agent stuffs entire modules into prompt semantic_search returns the 5 relevant paragraphs
Agent hallucinates the dependency chain trace_flow returns the verified path

End-to-End Workflow

Not just ingestion β€” a complete pipeline from source to agent query:

  1. Ingest β€” dev sync code / dev sync doc extracts structure and semantics; doc sync also generates dynamic YAKE entity rules per document before GLiNER extraction
  2. Store β€” graph relationships + vector embeddings persisted to disk
  3. Serve β€” MCP servers expose 30+ query tools on localhost
  4. Query β€” AI agents call MCP tools instead of reading raw files
  5. Update β€” next sync only processes changes, graph stays current

ADLC β€” Skills That Teach Agents How to Work

The ADLC skill pack (make install-adlc) installs 12+ battle-tested skills into your AI agent:

  • Plan before code β€” architecture-first planning, not dive-and-pray
  • Debug systematically β€” root cause analysis with evidence, not guess-and-check
  • Predict risks β€” 5-persona panel reviews changes before implementation
  • Generate edge cases β€” 12-dimension scenario coverage
  • Audit security β€” STRIDE + OWASP with automated fix suggestions

Works with: Claude Code, Qwen Code, OpenCode, GitHub Copilot, Cursor, Continue, Antigravity.


Quick Start

Prerequisites

  • Python 3.12+
  • uv (fast Python package manager)
  • Node.js β‰₯ 18 (only for ADLC skill pack)
  • (optional) .NET SDK 8+ and JDK 17+ β€” needed only for full-fidelity C#/VB.NET (Roslyn) and VB6 (ANTLR) parsing. See Parser Environment Requirements.

Install & Sync

git clone https://github.com/baka3k/cortex-harness.git
cd cortex-harness

make install          # create .venv, install deps, register global `dev` CLI
make storage-init     # initialize persistent storage at ~/.cortext-harness/v1/
make doctor           # verify Qdrant + FalkorDB connectivity

dev init              # interactive wizard β€” configure project, backend, folders
dev sync code         # ingest source code β†’ graph + embeddings
dev sync doc          # ingest documentation β†’ knowledge graph + vectors
dev start             # MCP servers: code (:8788) + doc (:8789)

Point your AI agent at http://localhost:8788 (code) or http://localhost:8789 (doc). The agent gets full graph + vector query access through standard MCP protocol.


Supported Languages

Compiler-grade parsers β€” not regex, not naive AST. Each analyzer uses the right tool for the language:

Language Parser Technology What It Extracts
C# / VB.NET Roslyn workspace (semantic model) Full type resolution, method invocations, LINQ, async/await, ASP.NET overlays
VB6 / VBA ANTLR4 (whole-program parse) Call graphs, control flow, designer forms, control arrays
Java / Kotlin Tree-sitter + Spring/Struts/MyBatis overlays Calls, generics, inheritance, framework routes, DI wiring
C / C++ Clang-based Pointers, templates, macros, header resolution
Go, Rust, Swift Tree-sitter Calls, types, interfaces, concurrency patterns
TypeScript / JavaScript Tree-sitter Imports, exports, async, React/Vue components
Python AST + import graph Decorators, generators, type hints
COBOL Custom parser (copybook-aware) Paragraphs, COPY resolution, file I/O
Perl, Delphi, Dart Tree-sitter / custom Module structure, calls, inheritance
Android Mixed Java + Kotlin analyzer Cross-language calls, Activity lifecycle

Framework overlays: Spring, Struts, MyBatis, Flutter, ASP.NET Core, ASP.NET Framework, Express.js, FastAPI/Django, Laravel

Document formats: PDF, Markdown, DOCX, PPTX, XLSX, TXT

Parser Environment Requirements

Some analyzers depend on an external toolchain. Prepare the ones below to get full-fidelity parsing:

Environment Required By Behavior When Missing
.NET SDK 8+ (dotnet on PATH) C# / VB.NET parsers β€” Roslyn semantic engine Falls back to tree-sitter (C#) / regex (VB.NET) parsing with a logged warning
JDK 17+ (java on PATH) VB6/VBA parser β€” ANTLR engine (worker runs as a java -jar process) Engine auto falls back to the regex engine with a one-line warning; forcing engine=antlr fails loudly
Maven 3.9+ Rebuilding the VB6 ANTLR worker jar only Optional at parse time β€” a prebuilt jar ships with the repo
ANTLR4 VB6/VBA whole-program parse Nothing to install separately β€” the grammar is bundled inside the vendored worker jar and runs on the JDK above

Notes:

  • Roslyn workers self-build: the first C#/VB.NET run executes dotnet build against the bundled worker projects (code-tiny/tools/csharp/roslyn_worker/, code-tiny/tools/vb/roslyn_worker/); a matching prebuilt DLL is reused when available.
  • Java & Kotlin need no external toolchain: Java sources are parsed with tree-sitter grammars installed via pip (tree-sitter-java, tree-sitter-kotlin). Parsing Java code does not require a JDK β€” the JDK above is consumed only by the VB6 ANTLR engine.

Default run vs. standalone runs. The default sync (dev sync code / dev sync code all) scans every configured language analyzer, so a missing toolchain is reported as a warning or per-language error for exactly the parts that need it. When you run an individual analyzer standalone, you only need the toolchain for the languages that analyzer covers β€” e.g. syncing a Java-only repo requires none of the environments above.


Storage

Embedded by default β€” data lives on disk, no network required.

~/.cortext-harness/v1/instances/default/
β”œβ”€β”€ manifest.json
β”œβ”€β”€ qdrant/{code,doc}/              # vector embeddings
β”œβ”€β”€ falkordb/{code,doc}/data.rdb    # graph relationships
└── backups/

Remote mode for shared team infrastructure:

make infra-up     # start Docker: Qdrant (:6333) + FalkorDB (:6379)
dev init          # select "remote" backend, enter connection details

Multiple projects share one storage instance with isolated graph/collection namespaces. Override paths with CORTEX_DATA_HOME or CORTEX_STORAGE_INSTANCE.


CLI Reference

Lifecycle

Command Description
dev build Create/reuse .venv, install dependencies
dev doctor Connectivity probes + MCP port diagnostics
dev start Launch MCP servers (code + doc)
dev start --server code --name shop --port 8790 Named instance with custom port
dev stop Stop all MCP servers
dev stop --name shop Stop one named instance

Sync

Command Description
dev sync code Incremental code sync (graph + embeddings)
dev sync code all Full sync across all configured source roots
dev sync code --sync-mode graph --full-scan Graph-only rebuild
dev sync code --sync-mode embedding --full-scan Embedding-only rebuild
dev sync doc Incremental doc sync (GraphRAG + vectors)
dev sync doc all Full doc sync across all configured doc roots

Configuration & Maintenance

Command Description
dev init Interactive setup wizard
dev status Show active config
dev ignore add <FOLDER>... Add folders to scan-ignore list
dev install-adlc Bootstrap ADLC skill pack for AI agents
make export-db OUTPUT=project.cortexdb Export project database
make import-db ARCHIVE=project.cortexdb Import project database

MCP Tools for AI Agents

One server, every query pattern:

Category Key Tools
Symbol & graph search_functions, get_symbol, query_subgraph, explore_graph
Flow tracing trace_flow, reconstruct_flow, find_workflows_containing
Impact analysis analyze_workflow_impact, find_paths, compute_scc
Fullstack bridge get_api_call_chain, find_callers_of_endpoint
Semantic search semantic_search, search_by_code
Project context get_project_modules, get_framework_context, get_public_apis
Dependency planning topological_sort, plan_dependency_order

Architecture

cortex-harness/
β”œβ”€β”€ cortex_harness/     # CLI (dev.py), config, storage management
β”œβ”€β”€ code-tiny/          # Language analyzers, MCP servers, graph writers
β”‚   β”œβ”€β”€ tools/          # Per-language analyzers (csharp, java, go, rust, ...)
β”‚   β”œβ”€β”€ mcp/            # Unified MCP server + per-language backends
β”‚   └── tools/graph/    # Shared graph schema, writers, CLI helpers
β”œβ”€β”€ doc-tiny/           # Document pipeline, GraphRAG, entity extraction
└── scripts/            # Lifecycle, MCP management, migrations

Design principles:

  • Embedded-first β€” zero infrastructure for individual developers
  • Provider-neutral β€” swap FalkorDB ↔ Neo4j without touching analyzers
  • Project-scoped β€” multiple projects, one storage, isolated namespaces
  • Incremental by default β€” Git-aware, SHA-256 verified, only changed files
  • Crash-safe β€” write-ahead journal for graph mutations with automatic replay

Troubleshooting

Windows: ModuleNotFoundError: No module named 'requests'
uv pip install --python .venv\Scripts\python.exe --requirements code-tiny\requirements.txt
Windows: 'dev' is not recognized as a command

Run directly: .venv\Scripts\dev.exe <command>, or add %USERPROFILE%\.local\bin to PATH.

CUDA / PyTorch issues
uv pip install torch torchvision torchaudio --default-index https://download.pytorch.org/whl/cu128
python -c "import torch; print(torch.cuda.is_available())"
Port already in use

dev start --port 8790 to pick a different port, or dev stop to clear existing servers.


License

See LICENSE for details.

About

A cognition-aware context and harness orchestration framework that combines Graph DBs, Vector DBs, and LLMs to build structured, memory-consistent, and scalable AI systems.

Topics

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages