An industry-level, local-first document discovery, hybrid retrieval, and Retrieval-Augmented Generation (RAG) system built with Java 21 LTS, Spring Boot 3.x, Apache Lucene 9.x, Qdrant Dense Vector Search, Reciprocal Rank Fusion (RRF), and a React + TypeScript web interface.
- Multi-Format Ingestion: Recursively discovers and extracts structured text from PDF (page-aware), DOCX (paragraphs/headings), Markdown (heading hierarchies), HTML (sections), and TXT files.
-
Incremental Change Detection: Computes SHA-256 checksums to categorize files into
ADDED,MODIFIED(delete old chunks + re-index),DELETED(purge), andUNCHANGED(skip) states. - Natural Boundary Chunking: Preserves PDF pages, headings, and paragraph boundaries; merges small adjacent paragraphs up to target token budget; splits oversized sections with sentence boundaries; assigns deterministic SHA-256 chunk IDs.
-
Hybrid Retrieval:
-
Lexical: Apache Lucene inverted index with BM25 similarity, phrase boosts, and
<mark>snippet highlighting. - Semantic: Dense vector embedding (384d, Cosine similarity) with Qdrant and resilient embedded fallback.
-
Fusion: Reciprocal Rank Fusion (
$RRF(d) = \sum \frac{w_i}{k + \text{rank}_i(d)}$ ,$k=60$ ).
-
Lexical: Apache Lucene inverted index with BM25 similarity, phrase boosts, and
-
Context Engineering & Grounded RAG: Token-budgeted context construction, strict system prompt grounding, traceable citations (
[CIT-1]), and explicitINSUFFICIENT_EVIDENCEdetection. - Modern Web UI: React 18 + TypeScript + Tailwind CSS with live indexing progress, hybrid search with diagnostic score badges, expandable RAG evidence trays, and interactive benchmark evaluation metrics.
- Java 21 LTS
- Maven 3.9+ (or included Apache Maven)
- Node.js 18+ & npm
cd frontend
npm install
npm run build(The build outputs directly to src/main/resources/static so Spring Boot serves the UI automatically)
# Set JAVA_HOME to Java 21 if needed
$env:JAVA_HOME = "C:\Program Files\Java\jdk-21.0.10"
# Compile and run
mvn clean spring-boot:runNavigate to http://localhost:8080 in your web browser.
local-context-search/
├── pom.xml
├── README.md
├── docker-compose.yml
├── docs/
│ ├── INTERVIEW_PREPARATION_AND_SYSTEM_DESIGN.md # Master Interview & System Design Guide
│ ├── industry-standard-technical-documentation.md
│ ├── architecture.md
│ ├── api-specification.md
│ └── retrieval-evaluation.md
├── sample_corpus/
│ ├── Security_Design.md
│ ├── Hybrid_Search_Architecture.md
│ ├── PDF_Chunking_Specification.txt
│ ├── Vector_Database_Storage.html
│ └── Incremental_Ingestion_Engine.md
├── src/main/java/com/localcontext/
│ ├── LocalContextSearchApplication.java
│ ├── api/ # REST Controllers & DTOs
│ ├── ingestion/ # File Scanner, Extractor, Chunker, Incremental Indexer
│ ├── embedding/ # Dense Semantic Vector Embeddings
│ ├── lexical/ # Apache Lucene BM25 Search
│ ├── vector/ # Qdrant & Embedded Local Vector Store
│ ├── retrieval/ # Hybrid Search & RRF Fusion Engine
│ ├── rag/ # Context Builder, Prompt Builder, LLM Answering
│ ├── eval/ # Retrieval Evaluation Benchmark Suite
│ ├── model/ # Domain entities (DocumentChunk, Citation, SearchResult, FileRecord, IndexJob)
│ └── config/ # Configuration & Web MVC
├── src/main/resources/
│ ├── application.yml
│ └── static/ # Production build of React UI
├── src/test/java/com/localcontext/
│ ├── ingestion/
│ ├── retrieval/
│ └── rag/
└── frontend/
├── package.json
├── vite.config.ts
├── tailwind.config.js
└── src/
├── App.tsx
├── components/ # Navbar, SearchTab, AskTab, IndexTab, EvalTab, ChunkModal
├── services/ # API client
└── types/ # TypeScript models
mvn test