DHARA
An agentic RAG engine for Indian legal research. It searches Supreme Court judgments with hybrid retrieval, reranks twice, and runs a LangGraph agent that plans, drafts and checks its answer.
Overview
Legal questions are hard for plain RAG. Judgments run to tens of thousands of words, they mix the parties’ arguments with the court’s reasoning, and the details that matter are exact: a section number, a statute, a precedent.
DHARA makes the corpus legal-aware before anything is embedded, then lets an agent drive retrieval instead of searching once and hoping.
Making the corpus legal-aware
Before anything is embedded, every judgment goes through a legal NLP pipeline:
Clean
400 judgments of the Supreme Court of India are cleaned and normalized.
Label rhetorical roles
Each sentence is tagged with one of 13 rhetorical roles, such as facts, issues, arguments, analysis, ratio and the final ruling, so the court’s reasoning can be told apart from what the parties argued.
Extract legal entities
OpenNyAI’s legal NER model finds 14 entity types, including statutes, provisions, precedents, courts and judges. Post-processing clusters precedents and pairs each provision with its statute.
Chunk and index
The merged records become 33,236 passages, each carrying its rhetorical role and the legal entities, concepts and keywords it mentions, indexed into two Pinecone indexes.
Answering a question
A LangGraph state machine runs five nodes and can loop:
Analyze
Gemini 2.5 Flash classifies the question: the kind of research, the legal concepts, search terms and complexity.
Extract
The agent finds the three most relevant passages, opens the judgments they come from and reads their precedents and provisions to find the statutes that apply.
Retrieve and draft
Hybrid retrieval returns the top passages with relevance scores, and the model drafts an answer that must cite provisions and precedents from that context, or say the context is not enough.
Synthesize
Everything is merged into a five-part brief: summary, legal framework, case law, practical implications and next steps.
Check
A quality gate scores the brief and can send the agent back for another pass, up to three rounds.
Hybrid retrieval and two-stage reranking
A sparse index catches exact legal tokens like “Section 498A” or “Article 21”. A dense index built on Gemini embeddings (3,072 dimensions) catches meaning when the wording differs.
The sparse side returns 20 candidates that Pinecone’s hosted bge-reranker-v2-m3 trims to 10, the dense side returns 10, and a MiniLM cross-encoder reranks the merged set down to the final few passages.
Results
Each stage buys accuracy with a little latency:
| Configuration | Search accuracy | P95 latency | Top-10 precision | Recall@50 |
|---|---|---|---|---|
| Basic retrieval (no NER, no reranking) | 85% | 250ms | 78% | 82% |
| NER only | 89% | 320ms | 85% | 86% |
| Hybrid retrieval, no reranking | 91% | 380ms | 88% | 89% |
| Full pipeline: NER, hybrid, double reranking | 97% | 500ms | 95% | 94% |
From my evaluation, as reported in the repository.
Serving it
A FastAPI service loads the agent once at startup and exposes research, summarize and entity-extraction endpoints, with structured JSON logs for every request. The Docker image bakes in the cross-encoder so containers start without downloading models, API keys arrive as Docker secrets, and the service runs on AWS behind a load balancer.