Advanced local-first RAG system powered by Ollama and LangGraph. Optimized for high-performance sLLM orchestration featuring adaptive intent routing, semantic chunking, intelligent hybrid search (FAISS + BM25), and real-time thought streaming. Includes integrated PDF analysis and secure vector caching.
python nlp semantic-search reranking faiss rag fastapi streamlit vector-database hybrid-search langchain pdf-chat local-ai ollama semantic-chunking langgraph ai-orchestration sllm thought-streaming
-
Updated
Sep 1, 2026 - Python