A Korean legal question-answering prototype that asks targeted clarification questions before retrieval and generation when the user's request is underspecified.
Research prototype: outputs are for informational purposes and are not legal advice.
Legal questions that look similar on the surface can require different answers depending on employment type, timing, jurisdiction, workplace size, or other missing facts. ClarifyLaw treats ambiguity resolution as part of the retrieval pipeline rather than forcing the generator to answer an incomplete question.
User query
│
├─ Query and topic classification
│
├─ Ambiguity detection
│ └─ Topic-specific clarification and slot tracking
│
├─ Query rewrite
│
├─ Hybrid retrieval: BM25 + dense embeddings + reranking
│
└─ LLM answer generation with citations and a legal disclaimer
| Component | Implementation |
|---|---|
| Primary domain | Korean labor-law questions |
| Interaction policy | Clarify-then-Answer, up to two clarification rounds |
| Retrieval | BM25 + dense retrieval + optional reranker |
| Generation | Configurable Korean instruction/legal model integration |
| Interface | Gradio application and web-server entry points |
| Evaluation | Component tests plus scripts for retrieval and LLM evaluation |
- topic routing for common labor-law question families;
- topic-specific slot sets and dynamic clarification questions;
- conversation-state tracking and duplicate-question avoidance;
- rewritten queries that incorporate the user's clarification;
- hybrid retrieval with source metadata and optional reranking;
- answer post-processing and an informational-use disclaimer;
- mock-mode and component-level tests for development without a full GPU stack.
The repository contains multiple generations of experiment, poster, and deployment notes. Some historical documents mix measured results with expected or illustrative values. In particular:
- the claimed “71% ambiguity reduction” was not supported by the project audit;
- several answer-quality numbers were documented as anticipated LLM-evaluator outcomes rather than a completed independent evaluation;
- an evaluation loss value alone does not establish absence of overfitting.
Accordingly, this landing page makes implementation claims only. Quantitative
results should be promoted after the associated inputs, evaluator outputs, and
aggregation script are packaged into one reproducible artifact. See
CONCLUSION_FACT_CHECK.md for the historical
claim audit.
git clone https://github.com/cosmic4dev/legal_bot.git
cd legal_bot
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python app.pyThe default interface is served locally by Gradio. Model access, data paths,
and retrieval options are configured in configs/model_config.yaml.
python scripts/run_tests.pySome tests require model downloads or locally prepared legal data. Start with the clarification and slot-tracking unit tests when running without a GPU.
src/ Core classification, clarification, retrieval, and generation
configs/ Model and pipeline configuration
prompts/ Clarification and answer prompt templates
scripts/ Data preparation, training, evaluation, and setup utilities
test_*.py Component and integration test entry points
docs/ Architecture and research documentation
This public snapshot intentionally focuses on the core prototype and selected technical documentation rather than historical deployment and poster material.
- Large datasets, vector stores, and fine-tuned checkpoints are not intended to be committed to Git.
- AI Hub and other external data must be obtained under their respective terms.
- API and Hugging Face credentials must be supplied through environment variables or local secret stores.
- Before production use, rebuild the retrieval index from an audited legal corpus and validate every citation path.
- The current focus is Korean labor law; behavior outside that scope is not established.
- Clarification quality depends on the topic taxonomy and slot coverage.
- Retrieved passages and generated citations require independent legal verification.
- Historical scripts contain several environment-specific deployment paths.
- The repository currently has no repository-level license, so no blanket open-source permission is implied.
See PUBLIC_RELEASE.md for the snapshot boundary and
provenance.
Contributions are especially useful in:
- ambiguity benchmarks with gold clarification questions;
- end-to-end evaluation that separates retrieval, clarification, and generation;
- citation-faithfulness testing against versioned statutes;
- portable CPU/mock-mode demos and dependency cleanup;
- user studies comparing immediate answers with clarification-first interaction.