Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

⚖️ LexBridge

RAG-based legal document simplifier — upload contracts, agreements, or legal documents and ask questions in plain English. LexBridge uses Retrieval-Augmented Generation to break down dense legal jargon into answers anyone can understand.


🧩 Problem

Legal documents — rental agreements, terms of service, contracts — are written in dense, jargon-heavy language that's hard for non-lawyers to parse. Most people sign things they don't fully understand because reading a 20-page lease or ToS document line-by-line isn't realistic.

LexBridge lets you upload a document once and ask it direct questions like "What happens if I miss a payment?" or "Can I terminate this early?" — getting plain-language answers grounded in the actual document, with source excerpts so you can verify the answer yourself.


⚙️ How It Works

  1. Upload — User uploads a PDF or TXT legal document via the Streamlit interface.
  2. Chunk — The document is split into overlapping text chunks using RecursiveCharacterTextSplitter to preserve context across boundaries.
  3. Embed — Each chunk is converted into a vector embedding using sentence-transformers/all-MiniLM-L6-v2.
  4. Store — Embeddings are stored in a local ChromaDB vector store.
  5. Retrieve — When a user asks a question, the top-k most relevant chunks are retrieved via similarity search.
  6. Generate — The retrieved context is passed to Groq's LLaMA 3.1 model through a custom prompt that instructs it to explain legal concepts in plain language, and the answer is streamed back with source citations.

This is a standard Retrieval-Augmented Generation (RAG) pipeline, ensuring answers are grounded in the actual uploaded document rather than the model's general knowledge — reducing hallucination risk on document-specific questions like clause numbers or specific terms.


🛠️ Tech Stack

Layer Technology
Frontend / UI Streamlit
Orchestration LangChain
LLM Groq (LLaMA 3.1 8B Instant)
Embeddings HuggingFace sentence-transformers
Vector Store ChromaDB
Document Parsing PyPDF, Unstructured

📂 Project Structure

lexbridge/
├── app.py                 # Streamlit UI and chat interface
├── rag_engine.py           # Embedding, vector store, and RAG chain logic
├── document_loader.py      # PDF/TXT loading and chunking
├── requirements.txt
├── .env                     # API keys (not committed)
└── data/                    # Uploaded documents (runtime only)

🚀 Setup & Run Locally

# Clone the repo
git clone https://github.com/NeuralDarsh/LexBridge.git
cd LexBridge

# Create and activate a virtual environment
python -m venv venv
venv\Scripts\activate        # Windows
# source venv/bin/activate   # Mac/Linux

# Install dependencies
pip install -r requirements.txt

# Add your Groq API key
echo "GROQ_API_KEY=your_key_here" > .env

# Run the app
streamlit run app.py

Get a free Groq API key at console.groq.com.


🗺️ Roadmap

  • Retrieval/answer quality evaluation suite
  • Multi-document comparison (e.g. compare two contract versions)
  • Multilingual support, including Japanese legal documents
  • FastAPI backend + dedicated frontend
  • Dockerized deployment

👤 Author

Built by NeuralDarsh — AI/ML student exploring applied RAG systems and legal-tech use cases.

About

RAG-based legal document simplifier — upload contracts and get plain-English answers using LangChain, ChromaDB, and Groq's LLM.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages