Interruptible Reasoning at Thought Units (IRTU)
A final-output check can hide a bad answer, but it cannot show which premise needs repair before that premise steers a tool, a handoff, or the next chat turn. InterrupThink lets a supervisor interrupt a specialist at checkable ThoughtUnit boundaries and correct the reasoning path — not only cancel it, and not at raw tokens.
This is an experimental library, not a product or a hosted service.
Getting started · Examples · Cases · Concepts · Results · Live evaluation · Contributing · License
Needs Python 3.11+. A live run calls an OpenAI Responses-compatible endpoint.
Copy .env.example to .env and fill the model, base URL, and key.
cp .env.example .env
uv pip install -e .
python3 cases/freeze-write/run.pyfrom interrupthink import LiveLlm, LlmMonitor, run_session
result = run_session(
llm=LiveLlm(user_prompt="Check the claim, then answer."),
monitor=LlmMonitor(),
)
# result.committed_answer, interrupt_ids, dropped_ids, request_count, tool_callsLiveLlm reads LLM_MODEL, LLM_BASE_URL, and LLM_API_KEY (or
OPENAI_API_KEY). The default base URL is https://api.openai.com/v1.
Point LLM_BASE_URL at another endpoint that accepts the same request
shape. LlmMonitor reads the SUPERVISOR_* variables.
The same floor accepts a scripted specialist. This path needs no API key and returns the same result on every run.
python3 examples/dummy/run_session_dummy.pyfrom interrupthink import DummyTool, FakeLlm, ScriptedMonitor, run_session
tool = DummyTool()
llm = FakeLlm([wrong_xml, stopped_xml])
monitor = ScriptedMonitor(trigger_kind="premise", trigger_contains="already approved")
result = run_session(llm=llm, monitor=monitor, tool=tool)Host policy is a separate check. examples/dummy/tool_policy_deny.py refuses publish after the monitor returns Ok.
After an interrupt, FakeLlm must supply a second XML document or the session reports an error. run_session does not add application-specific prompts to the request.
Local wheel (still not PyPI):
uv build --wheel
uv pip install --offline --no-index dist/interrupthink-0.0.1-py3-none-any.whlLanguage models can produce a convincing answer even when it rests on stale or missing information. That matters more when an answer does not end the work. In a chat, a wrong claim can become history that shapes the next turn. In a workflow, it can lead to a proposed tool action or be passed to another specialist.
Imagine a release assistant that reads an old note saying a release is approved. It may prepare to publish, then report that publishing is safe. A final-output review can still reject the visible answer. However, by itself it does not give the application a clear checkpoint for the approval claim, the planned action, and the work that was already sound. The usual fallback is to discard the whole response and start again.
InterrupThink adds those checkpoints before the application commits an answer
or action. The application emits small, structured task steps; a supervisor
checks them; and the application can either stop an unsupported path or supply
the missing fact and continue from the last accepted step. The supervisor is
not a truth oracle: it works from the evidence and policy the application
provides, and Unknown remains the default when that evidence is insufficient.
A further reason is a long task in one process. A policy can stay written in the context while the specialist's own reasoning accumulates. Public reports of instruction drift describe the same pattern: the model follows the trajectory it has already produced, and the earlier policy no longer shapes the next step. InterrupThink is built so a monitor outside the model can read each semantic step. When that path no longer matches the policy, the run stops and continues from the last accepted step. The checks in this repository do not yet show that contrast on a long task.
The host application keeps its own graph, roles, or tools. run_session is
the checkpoint boundary: the specialist emits checkable ThoughtUnit steps,
the supervisor monitors them, and a tool or answer is committed only when
allowed. A normal final-output gate remains useful too; this floor gives it
earlier, structured context.
This makes two responses possible. The floor can block an unsafe action or
unsupported answer. Or it can correct the path: drop the rejected tail,
add a Patch, and resume from the last accepted step instead of restarting
the whole task. This is especially useful when the host would otherwise carry
a wrong claim into a later chat turn, a tool call, or a named handoff.
“Reasoning” here means structured task steps the application chooses to emit, such as a plan, premise, claim, or tool intent. It is not a mechanism for reading or transporting hidden chain-of-thought.
The default after a cut is rollback: inject a correction, drop the invalid tail, and continue from the last accepted step. A full restart is the fallback. The monitor default is Unknown (do not cut). Import the package as interrupthink.
flowchart TB
host[Host application]
session["run_session"]
specialist[Specialist]
units[ThoughtUnit steps]
monitor[Supervisor monitor]
commit[Commit tool or answer]
correct[Inject Patch]
block[Block unsafe action]
resume[Rollback with watermark]
host --> session
session --> specialist
specialist --> units
units --> monitor
monitor -->|Ok| commit
monitor -->|Unknown| hold[Hold answer / continue steps]
monitor -->|Patch| correct
monitor -->|False| block
correct --> resume
block --> resume
resume --> specialist
commit --> host
| Piece | Role |
|---|---|
ThoughtUnit |
One checkable step in the specialist's document |
run_session |
Thinking floor: the specialist works, the monitor observes, and tools run only if allowed |
LlmMonitor / ScriptedMonitor |
On-demand blocking verdict per ThoughtUnit; default Unknown |
Patch / False |
Interrupt: block a bad path, or inject a correction |
| rollback | Resume mid-stream from a checkpoint — not restart from scratch |
Escalation |
Package for the next specialist. The first specialist does not resume |
Consult |
Package for the checker. A Patch returns to the same specialist |
Takeover |
Package for an editor or a human. A human does not start a second specialist |
Wire the floor yourself. There is no canned product helper required for the public path.
An Ok verdict can also name a receiver. Unknown and a blank name do not.
Escalation, Consult, and Takeover compose the package. They do not call
run_session. The host opens the next session only when the name is present.
The package keeps the original task, the receiver, the supervisor reason, the
kept steps, recorded tool results, and the calls that must not be repeated.
The unfinished answer is not committed.
More detail: docs/concepts.md.
The next specialist receives the package. The first specialist does not resume.
from interrupthink import Escalation, ScriptedMonitor, run_session
result = run_session(
llm=dealer,
monitor=ScriptedMonitor(
trigger_kind="claim",
trigger_contains="ask policy",
escalate_to="policy",
),
)
if result.escalate_to:
package = Escalation.from_result(task, result)
# Pass package.text() into a new run_session for the policy specialist.Call site: examples/dummy/supervisor_escalation.py.
Live demo: examples/langgraph_offer_escalation.py.
Cookbook: cases/langgraph-offer/.
The checker reads the package. The checker's answer comes back to the same specialist as a Patch.
from interrupthink import Consult, Patch, ScriptedMonitor, run_session
result = run_session(
llm=assistant,
monitor=ScriptedMonitor(
trigger_kind="claim",
trigger_contains="ask the checker",
consult_to="checker",
),
)
if result.consult_to:
package = Consult.from_result(task, result)
checker = run_session(llm=checker_llm, monitor=monitor)
note = next(event.payload for event in result.events if event.type == "floor.consult")
answer = checker.committed_answer or ""
patch = Patch(
from_agent="A",
target_unit_id=str(note["unit_id"]),
rollback_to=None,
diagnosis="consult",
missing=answer,
directive=answer,
)
continued = run_session(llm=assistant, monitor=monitor, resume_patch=patch)checker_llm should see package.text(). assistant is the same specialist as the first session.
Call site: examples/dummy/supervisor_consult.py.
Live demos: examples/langchain_rule_consult.py, examples/llamaindex_page_consult.py. Cookbooks: cases/langchain-rule/, cases/llamaindex-page/.
editor is another specialist and may receive a new run_session. human is not: return the package and do not start a second specialist.
from interrupthink import ScriptedMonitor, Takeover, run_session
result = run_session(
llm=support,
monitor=ScriptedMonitor(
trigger_kind="claim",
trigger_contains="the supervisor should take this",
takeover_to="human",
),
)
if result.takeover_to == "human":
package = Takeover.from_result(task, result)
# Return package.text() to the person. Do not open another run_session.Call site: examples/dummy/supervisor_takeover.py.
Live demo: examples/autogen_support_takeover.py.
Editor shape: examples/database_takeover.py.
Cookbook: cases/autogen-support/.
Short call sites after pip install -e .. Full stories live in cases/.
| File | Calls |
|---|---|
examples/dummy/run_session_dummy.py |
run_session, no API key |
examples/dummy/staging_migrate.py |
staging-migrate example, no API key (local helper) |
examples/dummy/freeze_push_dummy.py |
freeze + dummy push, no API key |
examples/dummy/host_loop_dummy.py |
host loop; HITL at the tool boundary, no API key |
examples/dummy/tool_policy_deny.py |
host denies publish after monitor Ok |
examples/dummy/supervisor_escalation.py |
Escalation: next specialist receives the package |
examples/dummy/supervisor_consult.py |
Consult: patch returns to the same specialist |
examples/dummy/supervisor_takeover.py |
Takeover: editor or human receives the package |
examples/langgraph_offer_escalation.py |
LangGraph escalation |
examples/langchain_rule_consult.py |
LangChain consultation |
examples/crewai_order_escalation.py |
CrewAI escalation |
examples/autogen_support_takeover.py |
AutoGen takeover to a human |
examples/llamaindex_page_consult.py |
LlamaIndex consultation |
Framework demos install the framework in the venv with pip, not as a core dependency. The full list is examples/README.md.
Case runners use LiveLlm and LlmMonitor with a personal .env. Do not commit keys. Deterministic, no-key call sites live in examples/dummy/.
The host framework provides application structure, not the semantic thinking loop. Each specialist slot still calls run_session. Optional integrations are installed separately and skipped safely when unavailable.
| Host | Cookbook | Think slot |
|---|---|---|
| Native | cases/two-specialists/ |
two run_session |
| LangGraph | cases/langgraph-pipe/ |
one node or two nodes, each run_session |
| LangChain | cases/langchain-specialist/ |
keep the thinking floor around the framework |
| LlamaIndex | cases/llamaindex-retrieve/ |
use the retriever for retrieval |
| CrewAI | cases/crewai-pipe/ |
keep run_session inside each role |
| AutoGen | cases/autogen-pipe/ |
keep run_session inside each agent |
Named handoffs use the same floor. The framework keeps its nodes, roles, or index. The host composes the package and opens the next step.
| Route | Host | Cookbook | Demo |
|---|---|---|---|
| Escalation | LangGraph | cases/langgraph-offer/ |
examples/langgraph_offer_escalation.py |
| Consultation | LangChain | cases/langchain-rule/ |
examples/langchain_rule_consult.py |
| Escalation | CrewAI | cases/crewai-order/ |
examples/crewai_order_escalation.py |
Takeover to human |
AutoGen | cases/autogen-support/ |
examples/autogen_support_takeover.py |
| Consultation | LlamaIndex | cases/llamaindex-page/ |
examples/llamaindex_page_consult.py |
See CONTRIBUTING.md. Evidence, limitations, and how to
extend the public docs: docs/.
uv pip install -e ".[dev]"
python3 -m pytest tests/test_public_api.py tests/test_library_packaging.py -x --tb=short -qThe library is experimental. It is not a product or a
multi-agent mesh. Public claims are limited to the checks in
docs/results.md (deterministic tests) and
docs/live-evaluation.md (opt-in live-model
checks).
Copyright 2026 BillyBSig. Licensed under the Apache License, Version 2.0.