Ansh SinghalAI/ML & backend engineer
Engineering write-up

Agentic RAG architecture: how I built DHARA, a legal research agent

A walkthrough of a working agentic RAG system over Supreme Court of India judgments: legal-aware preprocessing, hybrid dense and sparse retrieval, two-stage reranking, and a LangGraph loop that analyzes, retrieves, drafts and checks.

What is agentic RAG?

Retrieval-augmented generation (RAG) grounds a language model in your own documents: retrieve the passages that look relevant, put them in the prompt, generate an answer. Plain RAG does this once, in a straight line.

Agentic RAG is retrieval-augmented generation where an agent controls the retrieval. Instead of one search followed by one answer, the system analyzes the question, decides what to look up, calls tools, and checks its own draft before it responds. Retrieval becomes something the agent does, possibly more than once, rather than a fixed step bolted on the front.

That difference matters most when questions are complex, the corpus is messy and a wrong answer is expensive. Legal research is all three, which is why I built DHARA, an agentic RAG engine for Indian legal research.

DHARA works over 400 judgments of the Supreme Court of India. Three things make them hard to search:

  • They are long. The first judgment in the corpus alone runs past 100,000 characters.
  • They mix voices. A judgment holds the facts, what each side argued, what the lower courts decided and, finally, the court's own reasoning. A passage that sounds authoritative may be an argument the court rejected.
  • The details are exact. Lawyers ask about “Section 498A IPC”, a specific article or a named precedent. Dense embeddings are good at meaning but blur exact tokens like section numbers.

So the system has to know what kind of text each passage is, match exact legal references, and show where every claim came from.

The architecture at a glance

Offline: build the index
  1. 400 judgments
  2. Clean
  3. Rhetorical roles
  4. Legal NER
  5. Merge and chunk
  6. 33,236 passages
  7. Sparse + dense indexes
Online: answer a question
  1. Question
  2. Analyze
  3. Extract entities
  4. Retrieve and draft
  5. Synthesize
  6. Quality check
  7. Answer
Inside every retrieval
  1. Sparse top 20
  2. bge-reranker-v2-m3 keeps 10
  3. Dense top 10
  4. Merge
  5. MiniLM cross-encoder
  6. Top passages
DHARA in two halves: an offline pipeline that makes the corpus legal-aware, and an online LangGraph agent that searches it through hybrid, twice-reranked retrieval. The quality check can loop back for up to three rounds.

Offline, a legal NLP pipeline turns raw judgments into passages that carry structure. Online, a LangGraph agent answers questions by calling tools that search those passages. A FastAPI service in Docker serves the whole thing.

Before anything is embedded, every judgment goes through three passes.

Rhetorical roles

Each sentence gets one of 13 rhetorical role labels from OpenNyAI's scheme for Indian judgments: preamble, facts, issues, the arguments of each side, analysis, statutes, precedents relied on and not relied on, the ratio of the decision, rulings by the lower court and by the present court, and none. Here is how the indexed passages break down:

Rhetorical rolePassages
Analysis8,985
Precedents relied on3,953
Petitioner’s arguments3,945
Facts2,917
Statutes2,217
Preamble2,178
Respondent’s arguments782
Ruling by a lower court666
Ruling by the present court650
Issues435
Ratio of the decision322

These labels let the system tell the court's reasoning apart from a party's argument, which no embedding can do on its own. Analysis passages dominate, but the 322 passages that state the ratio are the ones that carry a judgment's binding reasoning.

Legal named entities

OpenNyAI's transformer-based legal NER model for spaCy (en_legal_ner_trf) tags 14 entity types: provisions, statutes, precedents, courts, judges, case numbers, petitioners, respondents, lawyers, witnesses, dates, organizations, places and other people. Its post-processing clusters different mentions of the same precedent and pairs each provision with the statute it belongs to. Provisions and statutes are the most common types: 15,221 passages mention a provision and 12,250 mention a statute.

Merge, chunk and index

A merge script combines the role labels, the entities and the cleaned text of each judgment into one record. Those records are chunked into 33,236 passages from 390 judgments, with a median length of about 730 characters. Every passage keeps its metadata. Here is a real one, trimmed:

{
  "primary_role": "ANALYSIS",
  "entity_types": ["COURT", "ORG", "PROVISION", "STATUTE"],
  "legal_concepts": ["Section 41(b)", "Specific Relief Act", "41(b)", "IBC"],
  "keywords": ["respondent", "held", "court", "plaintiff", "bench", "2009 cal 231"],
  "text_length": 750
}

Step 2: Hybrid retrieval with two indexes

DHARA keeps two Pinecone indexes over the same passages:

  • A sparse index built with Pinecone's hosted sparse embeddings. Sparse vectors behave like keyword search, so exact tokens such as “498A” or “Article 21” score highly.
  • A dense index of 3,072-dimensional embeddings from Google's gemini-embedding-001. Dense vectors capture meaning, so a question can match a passage that never uses its exact words.

The retriever's configuration, straight from the repo:

@dataclass
class RetrieverConfig:
    sparse_index_name: str = "legal-cases-pincone-sparse-2048"
    dense_index_name: str = "legal-cases-gemini-3072"
    sparse_top_k: int = 20
    dense_top_k: int = 10
    final_top_k: int = 5
    bert_rerank_model: str = "cross-encoder/ms-marco-MiniLM-L12-v2"
    pinecone_rerank_model: str = "bge-reranker-v2-m3"
    pinecone_rerank_top_k: int = 10

The sparse query asks for 20 candidates and the dense query for 10. Results are merged and deduplicated by passage ID. Keeping the indexes separate means each side can use its own embedding model, and the sparse side can rerank inside the same Pinecone call.

Step 3: Two-stage reranking

Retrieval is tuned for recall. Reranking turns that recall into precision, in two stages:

  1. Inside Pinecone. The sparse search includes a rerank step with Pinecone's hosted bge-reranker-v2-m3, which keeps the best 10 of the 20 sparse candidates.
  2. A cross-encoder over everything. After the merge, a MiniLM cross-encoder (ms-marco-MiniLM-L12-v2) reads each question and passage together and scores the pair. Only the top few survive.
# Condensed from app/core/custom_retriever.py
def rerank_with_bert(self, query, candidates, top_k):
    pairs = [[query, c["metadata"].get("text", "")] for c in candidates]
    scores = self.bert_reranker.predict(pairs)
    for c, score in zip(candidates, scores):
        c["bert_score"] = score
    return sorted(candidates, key=lambda c: c["bert_score"], reverse=True)[:top_k]

Embedding models encode the question and the passage separately, which is fast but blind to how they interact. A cross-encoder sees both at once, which is slower but much sharper, so it only runs on a short list. The Docker image downloads the cross-encoder at build time, so a fresh container never waits on a model download.

Step 4: The agent loop in LangGraph

The agent is a LangGraph StateGraph. Its state is a typed dictionary that every node reads and writes:

class LegalAgentState(TypedDict):
    messages: List[Any]
    original_query: str
    query_analysis: str
    extracted_entities: str
    basic_answer: str
    final_analysis: str
    confidence_score: float
    iteration_count: int
    needs_more_analysis: bool

The graph itself is five nodes and one conditional edge:

workflow = StateGraph(LegalAgentState)
workflow.add_node("analyze_query", self.analyze_query)
workflow.add_node("extract_entities", self.extract_legal_entities)
workflow.add_node("get_basic_answer", self.get_basic_legal_answer)
workflow.add_node("synthesize_analysis", self.synthesize_final_analysis)
workflow.add_node("quality_check", self.quality_check)

workflow.set_entry_point("analyze_query")
workflow.add_edge("analyze_query", "extract_entities")
workflow.add_edge("extract_entities", "get_basic_answer")
workflow.add_edge("get_basic_answer", "synthesize_analysis")
workflow.add_edge("synthesize_analysis", "quality_check")
workflow.add_conditional_edges(
    "quality_check",
    self.should_continue,
    {"continue": "extract_entities", "finish": END},
)

Every node runs on Gemini 2.5 Flash, at temperature 0.1 for orchestration and 0.2 inside the tools:

  1. Analyze the question. The model classifies the research needed (case law, statutory analysis or procedure), the legal concepts involved, the search terms to use and how complex the question is.
  2. Extract entities from similar judgments. The legal_entity_extractor tool retrieves the three most relevant passages, opens the judgments they come from and collects their precedent clusters and provisions. The model then decides which statutes, provisions and precedents apply. Because the entities were extracted offline, this costs one retrieval and one model call, not a read through whole judgments.
  3. Retrieve and draft. The legal_query_tool runs hybrid retrieval, keeps the top three passages and gives them to the model with their relevance scores. The prompt requires it to cite provisions and precedents from that context, and to say so plainly when the context is not enough.
  4. Synthesize. One more call merges the question analysis, the entities and the grounded draft into a five-part brief: executive summary, legal framework, case law analysis, practical implications and recommendations.
  5. Quality check. The brief gets a confidence score, and the conditional edge either finishes or sends the agent back to entity extraction, for at most three rounds.

A third tool, a document summarizer, turns any judgment into an eight-part summary (case details, facts, issues, provisions, the court's analysis, the decision, precedents cited and significance) and has its own endpoint.

One honest note: today the quality check is a simple heuristic built on length and legal signals, so most questions finish in one pass. It is the first thing I would make smarter.

Step 5: Serving it

  • A FastAPI app builds the agent once, in its lifespan hook at startup, so no request pays for loading models.
  • Endpoints under /api/v1 cover research, summarize, entity extraction and status, plus a /health check.
  • Middleware logs every request with structured fields, and each agent stage logs its own timings.
  • The Docker image is built on python:3.10-slim with uv and the reranker cached inside. API keys are mounted as Docker secrets, and Compose caps memory at 3 GB with a health check.
  • In production it runs on AWS ECS behind an Application Load Balancer, with TLS from AWS Certificate Manager.

Results

I measured each stage by switching it on in turn:

ConfigurationSearch accuracyP95 latencyTop-10 precisionRecall@50F1
Basic retrieval (no NER, no reranking)85%250ms78%82%0.80
NER only89%320ms85%86%0.85
Hybrid retrieval, no reranking91%380ms88%89%0.88
Full pipeline: NER, hybrid, double reranking97%500ms95%94%0.94

From my evaluation, as reported in the DHARA repository.

Every stage buys accuracy with some latency. Legal NER gave the biggest single jump, four points for about 70 ms. The two rerankers together added six points for about 170 ms, which is worth it when a wrong citation costs far more than a fraction of a second.

What I would change next

  • Query both indexes at once. The sparse and dense searches run one after the other today. Running them concurrently cuts retrieval latency without changing a single result.
  • Verify citations instead of only asking for them. Replace the heuristic quality check with a verifier that confirms every cited section and precedent appears in the retrieved passages, and loops back when one does not.
  • Use the rhetorical roles at query time. Every passage stores its role, but retrieval does not filter on it yet. Boosting ratio and analysis passages for “what did the court hold” questions, and arguments for “what did the petitioner claim”, is a cheap win.
  • A lawyer-labelled evaluation set. Relevance judgments from practitioners would make the accuracy numbers mean more than any automatic metric can.

The code is on GitHub under the MIT license, and the DHARA case study has the short version. If you are building a RAG system and want a second pair of eyes, or hiring for AI/ML, backend or GenAI work, my inbox is open.

Written by Ansh SinghalAI/ML and backend engineer in Greater Noida, Delhi NCR. I architected the AI Mesh Firewall at CyberUltron and lead ZeroDayShield. Open to AI/ML, backend, AI security and GenAI roles in Noida, Gurugram, Bengaluru or remote.DHARA case studyGitHubLinkedInEmail me

Let's build something secure.

I'm an AI/ML and backend engineer in Greater Noida, Delhi NCR, open to AI/ML, backend, AI security and GenAI roles in Noida, Gurugram, Bengaluru or remote.

Hire meDownload resumeBack to the portfolio