Skip to content

Repository files navigation

InterrupThink

Interruptible Reasoning at Thought Units (IRTU)

English · Bahasa Indonesia

A final-output check can hide a bad answer, but it cannot show which premise needs repair before that premise steers a tool, a handoff, or the next chat turn. InterrupThink lets a supervisor interrupt a specialist at checkable ThoughtUnit boundaries and correct the reasoning path — not only cancel it, and not at raw tokens.

This is an experimental library, not a product or a hosted service.

Getting started · Examples · Cases · Concepts · Results · Live evaluation · Contributing · License

Quickstart

Needs Python 3.11+. A live run calls an OpenAI Responses-compatible endpoint. Copy .env.example to .env and fill the model, base URL, and key.

cp .env.example .env
uv pip install -e .
python3 cases/freeze-write/run.py
from interrupthink import LiveLlm, LlmMonitor, run_session

result = run_session(
    llm=LiveLlm(user_prompt="Check the claim, then answer."),
    monitor=LlmMonitor(),
)
# result.committed_answer, interrupt_ids, dropped_ids, request_count, tool_calls

LiveLlm reads LLM_MODEL, LLM_BASE_URL, and LLM_API_KEY (or OPENAI_API_KEY). The default base URL is https://api.openai.com/v1. Point LLM_BASE_URL at another endpoint that accepts the same request shape. LlmMonitor reads the SUPERVISOR_* variables.

Without a key

The same floor accepts a scripted specialist. This path needs no API key and returns the same result on every run.

python3 examples/dummy/run_session_dummy.py
from interrupthink import DummyTool, FakeLlm, ScriptedMonitor, run_session

tool = DummyTool()
llm = FakeLlm([wrong_xml, stopped_xml])
monitor = ScriptedMonitor(trigger_kind="premise", trigger_contains="already approved")
result = run_session(llm=llm, monitor=monitor, tool=tool)

Host policy is a separate check. examples/dummy/tool_policy_deny.py refuses publish after the monitor returns Ok.

After an interrupt, FakeLlm must supply a second XML document or the session reports an error. run_session does not add application-specific prompts to the request.

Local wheel (still not PyPI):

uv build --wheel
uv pip install --offline --no-index dist/interrupthink-0.0.1-py3-none-any.whl

Why this exists

Language models can produce a convincing answer even when it rests on stale or missing information. That matters more when an answer does not end the work. In a chat, a wrong claim can become history that shapes the next turn. In a workflow, it can lead to a proposed tool action or be passed to another specialist.

Imagine a release assistant that reads an old note saying a release is approved. It may prepare to publish, then report that publishing is safe. A final-output review can still reject the visible answer. However, by itself it does not give the application a clear checkpoint for the approval claim, the planned action, and the work that was already sound. The usual fallback is to discard the whole response and start again.

InterrupThink adds those checkpoints before the application commits an answer or action. The application emits small, structured task steps; a supervisor checks them; and the application can either stop an unsupported path or supply the missing fact and continue from the last accepted step. The supervisor is not a truth oracle: it works from the evidence and policy the application provides, and Unknown remains the default when that evidence is insufficient.

A further reason is a long task in one process. A policy can stay written in the context while the specialist's own reasoning accumulates. Public reports of instruction drift describe the same pattern: the model follows the trajectory it has already produced, and the earlier policy no longer shapes the next step. InterrupThink is built so a monitor outside the model can read each semantic step. When that path no longer matches the policy, the run stops and continues from the last accepted step. The checks in this repository do not yet show that contrast on a long task.

How it works

Check decisions before committing them

The host application keeps its own graph, roles, or tools. run_session is the checkpoint boundary: the specialist emits checkable ThoughtUnit steps, the supervisor monitors them, and a tool or answer is committed only when allowed. A normal final-output gate remains useful too; this floor gives it earlier, structured context.

This makes two responses possible. The floor can block an unsafe action or unsupported answer. Or it can correct the path: drop the rejected tail, add a Patch, and resume from the last accepted step instead of restarting the whole task. This is especially useful when the host would otherwise carry a wrong claim into a later chat turn, a tool call, or a named handoff.

“Reasoning” here means structured task steps the application chooses to emit, such as a plan, premise, claim, or tool intent. It is not a mechanism for reading or transporting hidden chain-of-thought.

The default after a cut is rollback: inject a correction, drop the invalid tail, and continue from the last accepted step. A full restart is the fallback. The monitor default is Unknown (do not cut). Import the package as interrupthink.

flowchart TB
  host[Host application]
  session["run_session"]
  specialist[Specialist]
  units[ThoughtUnit steps]
  monitor[Supervisor monitor]
  commit[Commit tool or answer]
  correct[Inject Patch]
  block[Block unsafe action]
  resume[Rollback with watermark]
  host --> session
  session --> specialist
  specialist --> units
  units --> monitor
  monitor -->|Ok| commit
  monitor -->|Unknown| hold[Hold answer / continue steps]
  monitor -->|Patch| correct
  monitor -->|False| block
  correct --> resume
  block --> resume
  resume --> specialist
  commit --> host
Loading
Piece Role
ThoughtUnit One checkable step in the specialist's document
run_session Thinking floor: the specialist works, the monitor observes, and tools run only if allowed
LlmMonitor / ScriptedMonitor On-demand blocking verdict per ThoughtUnit; default Unknown
Patch / False Interrupt: block a bad path, or inject a correction
rollback Resume mid-stream from a checkpoint — not restart from scratch
Escalation Package for the next specialist. The first specialist does not resume
Consult Package for the checker. A Patch returns to the same specialist
Takeover Package for an editor or a human. A human does not start a second specialist

Wire the floor yourself. There is no canned product helper required for the public path.

An Ok verdict can also name a receiver. Unknown and a blank name do not. Escalation, Consult, and Takeover compose the package. They do not call run_session. The host opens the next session only when the name is present. The package keeps the original task, the receiver, the supervisor reason, the kept steps, recorded tool results, and the calls that must not be repeated. The unfinished answer is not committed.

More detail: docs/concepts.md.

Escalation

The next specialist receives the package. The first specialist does not resume.

from interrupthink import Escalation, ScriptedMonitor, run_session

result = run_session(
    llm=dealer,
    monitor=ScriptedMonitor(
        trigger_kind="claim",
        trigger_contains="ask policy",
        escalate_to="policy",
    ),
)
if result.escalate_to:
    package = Escalation.from_result(task, result)
    # Pass package.text() into a new run_session for the policy specialist.

Call site: examples/dummy/supervisor_escalation.py. Live demo: examples/langgraph_offer_escalation.py. Cookbook: cases/langgraph-offer/.

Consultation

The checker reads the package. The checker's answer comes back to the same specialist as a Patch.

from interrupthink import Consult, Patch, ScriptedMonitor, run_session

result = run_session(
    llm=assistant,
    monitor=ScriptedMonitor(
        trigger_kind="claim",
        trigger_contains="ask the checker",
        consult_to="checker",
    ),
)
if result.consult_to:
    package = Consult.from_result(task, result)
    checker = run_session(llm=checker_llm, monitor=monitor)
    note = next(event.payload for event in result.events if event.type == "floor.consult")
    answer = checker.committed_answer or ""
    patch = Patch(
        from_agent="A",
        target_unit_id=str(note["unit_id"]),
        rollback_to=None,
        diagnosis="consult",
        missing=answer,
        directive=answer,
    )
    continued = run_session(llm=assistant, monitor=monitor, resume_patch=patch)

checker_llm should see package.text(). assistant is the same specialist as the first session.

Call site: examples/dummy/supervisor_consult.py. Live demos: examples/langchain_rule_consult.py, examples/llamaindex_page_consult.py. Cookbooks: cases/langchain-rule/, cases/llamaindex-page/.

Takeover

editor is another specialist and may receive a new run_session. human is not: return the package and do not start a second specialist.

from interrupthink import ScriptedMonitor, Takeover, run_session

result = run_session(
    llm=support,
    monitor=ScriptedMonitor(
        trigger_kind="claim",
        trigger_contains="the supervisor should take this",
        takeover_to="human",
    ),
)
if result.takeover_to == "human":
    package = Takeover.from_result(task, result)
    # Return package.text() to the person. Do not open another run_session.

Call site: examples/dummy/supervisor_takeover.py. Live demo: examples/autogen_support_takeover.py. Editor shape: examples/database_takeover.py. Cookbook: cases/autogen-support/.

Examples

Short call sites after pip install -e .. Full stories live in cases/.

File Calls
examples/dummy/run_session_dummy.py run_session, no API key
examples/dummy/staging_migrate.py staging-migrate example, no API key (local helper)
examples/dummy/freeze_push_dummy.py freeze + dummy push, no API key
examples/dummy/host_loop_dummy.py host loop; HITL at the tool boundary, no API key
examples/dummy/tool_policy_deny.py host denies publish after monitor Ok
examples/dummy/supervisor_escalation.py Escalation: next specialist receives the package
examples/dummy/supervisor_consult.py Consult: patch returns to the same specialist
examples/dummy/supervisor_takeover.py Takeover: editor or human receives the package
examples/langgraph_offer_escalation.py LangGraph escalation
examples/langchain_rule_consult.py LangChain consultation
examples/crewai_order_escalation.py CrewAI escalation
examples/autogen_support_takeover.py AutoGen takeover to a human
examples/llamaindex_page_consult.py LlamaIndex consultation

Framework demos install the framework in the venv with pip, not as a core dependency. The full list is examples/README.md.

Case runners use LiveLlm and LlmMonitor with a personal .env. Do not commit keys. Deterministic, no-key call sites live in examples/dummy/.

Integrations

The host framework provides application structure, not the semantic thinking loop. Each specialist slot still calls run_session. Optional integrations are installed separately and skipped safely when unavailable.

Host Cookbook Think slot
Native cases/two-specialists/ two run_session
LangGraph cases/langgraph-pipe/ one node or two nodes, each run_session
LangChain cases/langchain-specialist/ keep the thinking floor around the framework
LlamaIndex cases/llamaindex-retrieve/ use the retriever for retrieval
CrewAI cases/crewai-pipe/ keep run_session inside each role
AutoGen cases/autogen-pipe/ keep run_session inside each agent

Named handoffs use the same floor. The framework keeps its nodes, roles, or index. The host composes the package and opens the next step.

Route Host Cookbook Demo
Escalation LangGraph cases/langgraph-offer/ examples/langgraph_offer_escalation.py
Consultation LangChain cases/langchain-rule/ examples/langchain_rule_consult.py
Escalation CrewAI cases/crewai-order/ examples/crewai_order_escalation.py
Takeover to human AutoGen cases/autogen-support/ examples/autogen_support_takeover.py
Consultation LlamaIndex cases/llamaindex-page/ examples/llamaindex_page_consult.py

Development

See CONTRIBUTING.md. Evidence, limitations, and how to extend the public docs: docs/.

uv pip install -e ".[dev]"
python3 -m pytest tests/test_public_api.py tests/test_library_packaging.py -x --tb=short -q

The library is experimental. It is not a product or a multi-agent mesh. Public claims are limited to the checks in docs/results.md (deterministic tests) and docs/live-evaluation.md (opt-in live-model checks).

License

Copyright 2026 BillyBSig. Licensed under the Apache License, Version 2.0.

About

Experimental Python library that interrupts specialist reasoning at semantic ThoughtUnit boundaries, corrects the path, and resumes from the last accepted step.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages