Ansh SinghalAI/ML & backend engineer
Agentic RAG · Legal research

DHARA

An agentic RAG engine for Indian legal research. It searches Supreme Court judgments with hybrid retrieval, reranks twice, and runs a LangGraph agent that plans, drafts and checks its answer.

Code on GitHubRead the architecture write-up
97%search accuracy
500msP95 response time
95%top-10 precision

Overview

Legal questions are hard for plain RAG. Judgments run to tens of thousands of words, they mix the parties’ arguments with the court’s reasoning, and the details that matter are exact: a section number, a statute, a precedent.

DHARA makes the corpus legal-aware before anything is embedded, then lets an agent drive retrieval instead of searching once and hoping.

Making the corpus legal-aware

Before anything is embedded, every judgment goes through a legal NLP pipeline:

  1. Clean

    400 judgments of the Supreme Court of India are cleaned and normalized.

  2. Label rhetorical roles

    Each sentence is tagged with one of 13 rhetorical roles, such as facts, issues, arguments, analysis, ratio and the final ruling, so the court’s reasoning can be told apart from what the parties argued.

  3. Extract legal entities

    OpenNyAI’s legal NER model finds 14 entity types, including statutes, provisions, precedents, courts and judges. Post-processing clusters precedents and pairs each provision with its statute.

  4. Chunk and index

    The merged records become 33,236 passages, each carrying its rhetorical role and the legal entities, concepts and keywords it mentions, indexed into two Pinecone indexes.

Answering a question

A LangGraph state machine runs five nodes and can loop:

  1. Analyze

    Gemini 2.5 Flash classifies the question: the kind of research, the legal concepts, search terms and complexity.

  2. Extract

    The agent finds the three most relevant passages, opens the judgments they come from and reads their precedents and provisions to find the statutes that apply.

  3. Retrieve and draft

    Hybrid retrieval returns the top passages with relevance scores, and the model drafts an answer that must cite provisions and precedents from that context, or say the context is not enough.

  4. Synthesize

    Everything is merged into a five-part brief: summary, legal framework, case law, practical implications and next steps.

  5. Check

    A quality gate scores the brief and can send the agent back for another pass, up to three rounds.

Hybrid retrieval and two-stage reranking

A sparse index catches exact legal tokens like “Section 498A” or “Article 21”. A dense index built on Gemini embeddings (3,072 dimensions) catches meaning when the wording differs.

The sparse side returns 20 candidates that Pinecone’s hosted bge-reranker-v2-m3 trims to 10, the dense side returns 10, and a MiniLM cross-encoder reranks the merged set down to the final few passages.

Results

Each stage buys accuracy with a little latency:

ConfigurationSearch accuracyP95 latencyTop-10 precisionRecall@50
Basic retrieval (no NER, no reranking)85%250ms78%82%
NER only89%320ms85%86%
Hybrid retrieval, no reranking91%380ms88%89%
Full pipeline: NER, hybrid, double reranking97%500ms95%94%

From my evaluation, as reported in the repository.

Serving it

A FastAPI service loads the agent once at startup and exposes research, summarize and entity-extraction endpoints, with structured JSON logs for every request. The Docker image bakes in the cross-encoder so containers start without downloading models, API keys arrive as Docker secrets, and the service runs on AWS behind a load balancer.

Deep diveAgentic RAG architecture: how I built DHARA, a legal research agentThe full architecture with code from the repo: preprocessing, hybrid retrieval, reranking, the LangGraph loop and what I'd change next.Read the write-up →

More projects

Let's build something secure.

I'm an AI/ML and backend engineer in Greater Noida, Delhi NCR, open to AI/ML, backend, AI security and GenAI roles in Noida, Gurugram, Bengaluru or remote.

Hire meDownload resumeBack to the portfolio