RAG-based legal document simplifier — upload contracts, agreements, or legal documents and ask questions in plain English. LexBridge uses Retrieval-Augmented Generation to break down dense legal jargon into answers anyone can understand.
Legal documents — rental agreements, terms of service, contracts — are written in dense, jargon-heavy language that's hard for non-lawyers to parse. Most people sign things they don't fully understand because reading a 20-page lease or ToS document line-by-line isn't realistic.
LexBridge lets you upload a document once and ask it direct questions like "What happens if I miss a payment?" or "Can I terminate this early?" — getting plain-language answers grounded in the actual document, with source excerpts so you can verify the answer yourself.
- Upload — User uploads a PDF or TXT legal document via the Streamlit interface.
- Chunk — The document is split into overlapping text chunks using
RecursiveCharacterTextSplitterto preserve context across boundaries. - Embed — Each chunk is converted into a vector embedding using
sentence-transformers/all-MiniLM-L6-v2. - Store — Embeddings are stored in a local ChromaDB vector store.
- Retrieve — When a user asks a question, the top-k most relevant chunks are retrieved via similarity search.
- Generate — The retrieved context is passed to Groq's LLaMA 3.1 model through a custom prompt that instructs it to explain legal concepts in plain language, and the answer is streamed back with source citations.
This is a standard Retrieval-Augmented Generation (RAG) pipeline, ensuring answers are grounded in the actual uploaded document rather than the model's general knowledge — reducing hallucination risk on document-specific questions like clause numbers or specific terms.
| Layer | Technology |
|---|---|
| Frontend / UI | Streamlit |
| Orchestration | LangChain |
| LLM | Groq (LLaMA 3.1 8B Instant) |
| Embeddings | HuggingFace sentence-transformers |
| Vector Store | ChromaDB |
| Document Parsing | PyPDF, Unstructured |
lexbridge/
├── app.py # Streamlit UI and chat interface
├── rag_engine.py # Embedding, vector store, and RAG chain logic
├── document_loader.py # PDF/TXT loading and chunking
├── requirements.txt
├── .env # API keys (not committed)
└── data/ # Uploaded documents (runtime only)
# Clone the repo
git clone https://github.com/NeuralDarsh/LexBridge.git
cd LexBridge
# Create and activate a virtual environment
python -m venv venv
venv\Scripts\activate # Windows
# source venv/bin/activate # Mac/Linux
# Install dependencies
pip install -r requirements.txt
# Add your Groq API key
echo "GROQ_API_KEY=your_key_here" > .env
# Run the app
streamlit run app.pyGet a free Groq API key at console.groq.com.
- Retrieval/answer quality evaluation suite
- Multi-document comparison (e.g. compare two contract versions)
- Multilingual support, including Japanese legal documents
- FastAPI backend + dedicated frontend
- Dockerized deployment
Built by NeuralDarsh — AI/ML student exploring applied RAG systems and legal-tech use cases.