A local-first, provider-neutral framework for adaptive voice conversations. It includes a campaign-driven conversation core, local Qwen LLM/ASR/TTS adapters, microphone and speaker interaction, a real-time call lifecycle, LiveKit room transport, outbound SIP, structured SQLite outcomes, operator tooling, and compliance-focused safeguards.
The distribution and repository use the general name adaptive-voice-agent. The
internal Python namespace remains speaking_agent for compatibility.
The complete local inference stack is currently verified on Apple-silicon macOS. See ROADMAP.md for the ordered cross-platform, CUDA, CPU, audio, packaging, benchmarking, and deployment TODOs.
A real PSTN call has not been placed from this repository because that requires private LiveKit credentials, an outbound trunk, and a number explicitly controlled by the operator. The command is implemented but deliberately blocked without all three.
Status last verified on 2026-08-20 on Apple-silicon macOS with Python 3.12:
| Area | Status |
|---|---|
| Campaign-driven conversation core | Complete and covered by scenario tests |
| Deterministic safety and field grounding | Complete for configured policies |
| Local Qwen LLM, ASR, and TTS | Verified with the pinned MLX models |
| Local microphone/speaker conversation | Verified in half-duplex and speaker modes |
| Local full duplex and barge-in | Implemented and tested; echo suppression remains experimental |
| LiveKit room audio | Verified bidirectionally |
| Outbound SIP/PSTN | Implemented behind controlled-call gates; real PSTN not yet exercised |
| SQLite outcomes, suppression, retention, and metrics | Complete and tested |
| DAMAC seller summaries, P1-P4 follow-ups, and lead webhook | Implemented and unit tested; external CRM not exercised |
| Approved market-data gateway | Implemented and unit tested; provider credentials/feed not supplied |
| Linux, Windows, CUDA, and portable inference | Not yet integrated or natively verified |
The current validation passes 256 tests, pip check, bytecode compilation, diff
validation, local Markdown-link checks, and editor diagnostics. A real local Qwen
planning smoke completed in 1.214 seconds for one warm single-turn scenario; this is a
smoke measurement, not a P50/P95 benchmark. Independent blocker review found no release
blockers for the current local prototype scope.
This is not yet a production telephony release. Production readiness still requires a controlled PSTN call, native platform evidence, measured latency and memory budgets, production acoustic echo cancellation/VAD, and authoritative legal/compliance review.
LiveKit / SIP -------> CallTransport adapter --------+
Incoming PCM -------> Turn detection |
v
Campaign JSON ---> ConversationPolicy ---> CallSession
| |
v v
ConversationState ASR / LLM / TTS protocols
| ^
v |
LeadOutcome mock / MLX adapters
|
+----> MarketDataProvider ---> approved API/feed gateway
|
v
CallRepository ---> SQLite ---> LeadWorkflowSink ---> CRM/Yasir
The domain, policy, conversation, and call lifecycle own application types. They do not import LiveKit, SIP, MLX, Qwen, SQLite, cloud SDKs, or OS-specific APIs. Provider types are translated only inside adapters.
The package requires Python 3.11+ and is currently verified on Python 3.12. Install the full verified local stack in the configured virtual environment:
.venv/bin/python -m pip install -e '.[realtime]'Pinned integration versions are mlx-lm==0.31.3, mlx-audio==0.4.8,
livekit-agents==1.6.10, and its required livekit==1.1.14.
Run the complete external-service-free test suite:
PYTHONPATH=src .venv/bin/python -m unittest discover -s tests -vRun the standard command-line quality gate before merging a behavioral change:
PYTHONPATH=src .venv/bin/python -m unittest discover -s tests -v
.venv/bin/python -m pip check
.venv/bin/python -m compileall -q src scripts
git diff --checkFor documentation changes, also verify local Markdown links and editor diagnostics with the available IDE/tooling. For platform or adapter changes, the native checks listed in the validation table below remain mandatory.
An editable install also provides the preferred public commands:
| Command | Purpose |
|---|---|
adaptive-voice-agent |
Text simulator with mock or MLX conversation model |
adaptive-voice-speech |
WAV-based ASR/TTS diagnostics |
adaptive-voice-local |
Local microphone/speaker conversation |
adaptive-voice-calls |
Inspect persisted outcomes and aggregate metrics |
adaptive-voice-retention |
Purge expired structured records |
The older speaking-agent-* command aliases and python -m speaking_agent... module
forms remain available for compatibility.
Run deterministically without models, speech, telephony, a database, or Internet:
PYTHONPATH=src .venv/bin/python -m speaking_agentUse the local Qwen model instead:
PYTHONPATH=src .venv/bin/python -m speaking_agent --model mlxThe default checkpoint is mlx-community/Qwen3-4B-Instruct-2507-4bit (about
2.26 GB). The simulator uses the same ConversationSession and campaign policy as a
voice call. Type /quit to end an unfinished simulation.
Generate PCM WAV audio with the macOS debug voice or Qwen3-TTS:
PYTHONPATH=src .venv/bin/python -m speaking_agent.speech_cli synthesize \
--adapter system --text "Local speech test" --output output/system.wav
PYTHONPATH=src .venv/bin/python -m speaking_agent.speech_cli synthesize \
--adapter qwen --voice Aiden --language English \
--text "I might sell in two months" --output output/qwen.wavTranscribe a 16-bit PCM WAV with Qwen3-ASR:
PYTHONPATH=src .venv/bin/python -m speaking_agent.speech_cli transcribe \
--language English --input output/qwen.wavThe harness reports TTS first-audio/total time and ASR first-partial/final time. The
default checkpoints are mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit
(about 1.97 GB) and mlx-community/Qwen3-ASR-0.6B-8bit (about 1.01 GB).
Talk to the complete local Qwen stack through the Mac microphone and speakers:
PYTHONPATH=src .venv/bin/python -m speaking_agent.local_voice_chatThe agent speaks the campaign opening, listens until 550 ms of trailing silence,
prints what Qwen3-ASR heard, runs the real conversation model, and speaks the Qwen3-TTS
response. Each turn reports ASR, LLM, TTS-first-audio, and speech-end-to-response
latency. The local helper applies a warm conversational speaking style by default; use
--style to replace it. Audio and transcripts are not written to disk by default.
Press Ctrl-C to stop.
On first use, allow microphone access for VS Code or the terminal in macOS System Settings > Privacy & Security > Microphone. Headphones prevent the microphone from hearing the agent's speaker output. List and select CoreAudio devices with:
PYTHONPATH=src .venv/bin/python -m speaking_agent.local_voice_chat --list-devices
PYTHONPATH=src .venv/bin/python -m speaking_agent.local_voice_chat \
--input-device 2 --output-device 3For telephone-like conversation, use full-duplex mode:
.venv/bin/adaptive-voice-local --full-duplexTo practice the personalized three-stage opening with fake DAMAC data:
.venv/bin/adaptive-voice-local --full-duplex --demo-metadataThat demo identifies the disclosed AI assistant, asks for Mr. Ahmed, mentions a fake
4-bedroom townhouse in Nice only after recipient confirmation, asks whether it is a good
time, and then starts qualification. NeoAI is the caller; Yasir is the human who receives
the post-call lead. When the two loopback demo services below are running,
--demo-metadata automatically reads fictional market data from port 8765, uses the
fixed fake number +971500000000, and sends Yasir's notification to port 8766 only
after the conversation reaches its closing outcome. The terminal then prints
Yasir notification sent. The configured campaign opening is used when metadata is
omitted.
Supply reviewed metadata explicitly when it becomes available:
.venv/bin/adaptive-voice-local --full-duplex \
--recipient-name "Ms. Fatima" \
--property-reference "your townhouse in Malta, DAMAC Lagoons" \
--phone-number "+971501234567" \
--project "DAMAC Lagoons" --cluster Malta --bedrooms 4 \
--property-type townhouseRecipient name and property reference must be supplied together. The reference gates the
personalized opening; the recipient name and verified structured fields may appear in the
retained call summary. Never populate metadata from an unverified enrichment source, and
apply the same access and deletion policy used for transcripts and call outcomes.
--phone-number must use E.164 format and is sent only to a configured lead workflow;
omit it when that workflow does not need a WhatsApp action.
Record a local full-duplex quality session only after obtaining and documenting the required consent:
.venv/bin/adaptive-voice-local --full-duplex --demo-metadata \
--record-audio \
--recording-consent-reference local-self-test-20260816The consent reference is an external audit identifier, not proof that consent was
legally sufficient. The operator remains responsible for approved notice, consent,
access, training use, and deletion rules. Local recording requires full duplex; without
--record-audio, no audio file is created.
Artifacts are written beneath data/recordings/YYYY-MM-DD/ by default:
<call-id>.wav: 16 kHz stereo PCM, owner on channel 1 (left) and agent on channel 2 (right). The agent channel accepts only transport-confirmed playout; queued audio discarded by barge-in is not recorded.<call-id>.json: consent reference, purpose, campaign/call IDs, timestamps, expiry, channel map, duration, and SHA-256 digest. It contains no transcript, phone number, recipient name, or property reference.
Directories are owner-only (0700) and files are owner-only (0600). Raw voice can
still identify a person and contain sensitive statements. The application does not
encrypt or de-identify recordings, label training examples, or establish rights to train
a model. Use encrypted storage and a separately reviewed de-identification/access/export
workflow before moving recordings off the controlled machine or using them for training.
Recording creation and retention refuse symlinked storage components, and an active-call
marker prevents retention from removing an in-progress call directory. Stale crash
partials/orphan WAVs are removed after the retention worker's grace period.
The microphone remains active while the agent speaks. Confirmed near-end speech stops the current TTS playback, preserves the interruption, and immediately enters the normal ASR/LLM/TTS flow. The terminal reports the transcript and interruption count without a separate “Listening” phase.
The personalized recipient-confirmation opening is a deliberate exception: it must finish before barge-in is enabled because it contains required disclosure and gates private property metadata. Audio spoken over that opening is queued; meaningful replies such as “Yes” are processed immediately after delivery, while low-information echo-like ASR fragments such as “Ah”, “Um”, “Hi”, or “Hello” are ignored. Normal barge-in resumes for all later turns. Unrecognized confirmation replies are bounded and cannot replay the opening indefinitely.
Terminal hangup also checks for microphone speech that began just before closing playout completed. A short bounded grace lets candidate speech cross the VAD threshold; confirmed speech is then allowed to finish so late do-not-contact requests can override the prior outcome and persist before hangup. Calls with no pending microphone activity still hang up immediately after playout.
Privacy-sensitive replies are also deferred from the first candidate speech frame, not only after VAD confirmation. This prevents an earlier queued “Yes” from causing property metadata playback while a later wrong-recipient correction is still being spoken.
With speakers, full-duplex mode uses the outgoing PCM as an echo reference and rejects
correlated or low-energy loopback. This is an experimental software echo suppressor, not
production-grade acoustic echo cancellation. Keep speaker volume moderate. If the agent
interrupts itself, increase --barge-in-energy-threshold or --echo-gain; if your voice
cannot interrupt it, lower those values. --echo-correlation-threshold controls how
closely microphone audio must match recent speaker output before it is rejected.
Use half-duplex speaker mode as the stable fallback in a difficult room:
.venv/bin/adaptive-voice-local --speaker-modeThe microphone stays closed while the agent speaks, waits 500 ms for room echo to
settle, and then starts listening automatically. Keep speaker volume moderate. In a very
echoey room, increase the guard with --speaker-settle-ms 1000. macOS default devices
are preferred because CoreAudio numeric indices change when monitors, docks, or headsets
connect. Use --list-devices and explicit device names only when defaults are wrong.
The local VAD accepts speech as short as 60 ms so one-word replies such as “yes”, “no”,
or “both” are retained; raise --minimum-speech-ms if brief room noise triggers it.
If normal speech is not detected, lower --energy-threshold from 0.02 to 0.01.
Raise it in a noisy room. Voice character can be tested with --voice and --style:
PYTHONPATH=src .venv/bin/python -m speaking_agent.local_voice_chat \
--voice Aiden --style "Speak warmly, naturally, and concisely."Qwen TTS now defaults to greedy decoding (--tts-temperature 0) so the same
text/voice/style produces repeatable PCM and is much less likely to drift in speaker
identity or delivery between turns. A native check on the pinned checkpoint produced
identical hashes for two repeated adapter generations. To deliberately restore more
prosodic variation, use for example --tts-temperature 0.3 --tts-top-k 30; sampled
decoding can also reintroduce audible voice/style variation.
Both modes use the same campaign, conversation policy, local models, and structured call state. LiveKit is the intended telephony transport path, where endpoint/WebRTC echo cancellation is expected to handle acoustic loopback.
Create .env from .env.example and provide your own values. At minimum, room tests
need LIVEKIT_URL, LIVEKIT_API_KEY, and LIVEKIT_API_SECRET.
Run the agent worker:
PYTHONPATH=src .venv/bin/python -m speaking_agent.livekit_worker devFor a local transport integration test, start the server in another terminal and run the bidirectional audio smoke script:
livekit-server --dev
PYTHONPATH=src .venv/bin/python scripts/livekit_audio_smoke.pyThe worker receives 16 kHz mono PCM, detects turns, supports barge-in by cancelling TTS and clearing the LiveKit audio queue, publishes 24 kHz mono PCM, waits for playout before hangup/transfer, and releases all session resources on exit.
The repository includes a loopback-only fake stack for development without DLD, portal,
CRM, WhatsApp, LiveKit, or PSTN credentials. Use only fake owners and controlled test
numbers. Every market response begins with FICTIONAL DEMO DATA, uses DEMO ONLY
sources, and is rejected automatically when controlled_test_mode is false.
Start the fictional market service:
PYTHONPATH=src .venv/bin/python scripts/demo_market_data_server.pyIt serves three fixtures: Nice 4BR townhouse, Malta 4BR townhouse, and Venice 6BR villa. Unknown combinations return no comparables instead of fabricated fallback prices.
In a second terminal, start the local Yasir/CRM receiver:
PYTHONPATH=src .venv/bin/python scripts/demo_lead_workflow_server.pyIt prints a compact Yasir notification, deduplicates by call_id, and appends the full
event to the owner-only ignored file data/demo/lead_events.jsonl. The latest received
event is available locally at http://127.0.0.1:8766/latest; the latest event that
actually requires a Yasir alert remains available at
http://127.0.0.1:8766/latest-yasir. Unqualified and not-interested calls are stored but
do not replace that alert view.
With both services running, start the real local Qwen voice conversation:
PYTHONPATH=src .venv/bin/python -m speaking_agent.local_voice_chat \
--full-duplex --demo-metadataAnswer the recipient and timing confirmations, complete the seller questions, and let
the agent deliver its closing line. The latest Yasir event changes only when the call is
complete; stopping with Ctrl-C before a final outcome does not create a lead.
For the identity question, say Yes or Yes, speaking; Okay is only an acknowledgement
and intentionally does not unlock the private property-reference step.
Alternatively, run a complete synthetic hot-seller flow without microphone or models:
PYTHONPATH=src .venv/bin/python scripts/demo_sales_flow.py \
--call-id demo-nice-hot-seller-1Change --call-id to create another event. To connect the running LiveKit worker to the
same fake services for a controlled call, configure:
SPEAKING_AGENT_MARKET_DATA_URL=http://127.0.0.1:8765/comparables
SPEAKING_AGENT_LEAD_WORKFLOW_URL=http://127.0.0.1:8766/events
The WhatsApp action is only a generated wa.me link and only appears after explicit
permission and number confirmation. The demo receiver never sends a WhatsApp message.
Do not expose either server beyond loopback or use its fictional evidence in a real owner
conversation.
Set SPEAKING_AGENT_MARKET_DATA_URL to an internal HTTPS gateway backed by approved DLD,
Property Monitor, REIDIN, or brokerage listing feeds. The application does not scrape
Property Finder or Bayut. It sends project, cluster, bedrooms, property_type, and
months as query parameters after the comparable property is identified. An optional
bearer token comes from SPEAKING_AGENT_MARKET_DATA_TOKEN.
The gateway response is normalized so completed sales cannot be confused with asking prices:
{
"actual_transactions": {
"source": "Dubai Land Department",
"as_of": "2026-08-20",
"count": 9,
"low_aed": 3550000,
"high_aed": 3950000,
"median_aed": 3750000
},
"current_listings": {
"source": "Approved brokerage feed",
"as_of": "2026-08-20",
"count": 12,
"low_aed": 3850000,
"high_aed": 4200000,
"median_aed": 4000000
},
"confidence": "high"
}Either evidence section may be omitted, but source and as_of are mandatory when it is
present. Remote endpoints must use HTTPS. Missing or failed market data never becomes a
model-generated estimate; qualification continues without the unavailable evidence.
Set SPEAKING_AGENT_LEAD_WORKFLOW_URL to an authorized HTTPS webhook. After each
connected human call, the worker first stores the call record and then posts the full
structured summary, transcript when enabled, recording URL when available, P1-P4
priority, follow-up timestamp/task flag, recommended action, and notification mode.
P1 hot sellers and P2 owners open at the right price set notify_yasir=true with
notification_mode=IMMEDIATE. P3 potential sellers use MARKET_FOLLOW_UP. P4 future
sellers create a dated follow-up task without an urgent notification. Qualified events
include an open_whatsapp_url only after the owner grants permission and confirms the
called number. The raw number is sent only to the
configured workflow endpoint and is not persisted in SQLite; stored records retain the
masked number. A failed webhook is logged and does not delete or roll back the call
record. Production deployments should place the endpoint behind a durable CRM/outbox
with authentication, retry, idempotency by call_id, and role-based access.
Outbound calling is restricted to controlled tests. Configure:
SIP_OUTBOUND_TRUNK_ID=ST_...
SPEAKING_AGENT_ALLOWED_TEST_NUMBERS=+<controlled-E.164-number>
SPEAKING_AGENT_SUPPRESSION_KEY=<at-least-32-random-bytes>
Keep the worker running, then inspect a masked dry run:
PYTHONPATH=src .venv/bin/python -m speaking_agent.call_cli +<controlled-number>Only after verifying the room, trunk, campaign identity, number, and calling policy:
PYTHONPATH=src .venv/bin/python -m speaking_agent.call_cli \
+<controlled-number> --executeOptional reviewed personalization uses the same flags as local mode and is forwarded in private LiveKit job metadata; the dry-run summary reports only whether personalization is enabled:
PYTHONPATH=src .venv/bin/python -m speaking_agent.call_cli \
+<controlled-number> \
--recipient-name "Mr. Ahmed" \
--property-reference "your townhouse in Nice, DAMAC Lagoons" \
--project "DAMAC Lagoons" --cluster Nice --bedrooms 4 \
--property-type townhouseFor a controlled LiveKit/SIP recording, first set behavior.recording_enabled to true
in the reviewed campaign, configure SPEAKING_AGENT_RECORDING_DIRECTORY, and dispatch
with a per-call consent reference:
PYTHONPATH=src .venv/bin/python -m speaking_agent.call_cli \
+<controlled-number> \
--recording-consent-reference consent-ticket-42 \
--executeThe worker blocks before dialing when recording is enabled but the consent reference is missing or invalid. A reference is required even if the campaign opening also contains a recording notice; configure both according to authoritative requirements.
The worker blocks non-allowlisted numbers, invalid E.164 input, active do-not-contact records, excessive attempts, too-short retry intervals, missing suppression keys, and unconfigured trunks. Non-test campaigns must additionally configure a reviewed timezone and local calling window. No jurisdiction-specific hours are invented by this example. The suppression key is bound to the SQLite database by a non-secret identifier; changing it fails closed. Key rotation requires an explicit migration, not a silent environment variable replacement.
Outbound sessions listen before speaking to separate a human from explicit voicemail or
IVR signatures. Human speech then receives the configured identity disclosure before the
adaptive response. Voicemail uses a separate campaign message. Optional cold transfer
uses SPEAKING_AGENT_TRANSFER_TO and requires provider support for SIP REFER.
When transfer is not configured, the agent uses the campaign's transfer-unavailable
message and hangs up safely instead of announcing a transfer it cannot perform.
Inspect recent structured call records:
PYTHONPATH=src .venv/bin/python -m speaking_agent.operator_cli list
PYTHONPATH=src .venv/bin/python -m speaking_agent.operator_cli show <call-id>
PYTHONPATH=src .venv/bin/python -m speaking_agent.operator_cli metricsSQLite defaults to data/speaking_agent.db. Records include the mandatory sectioned
seller summary, P1-P4 priority, follow-up timestamp, validated property/price/WhatsApp
fields, market evidence discussed, connection data, duration, and latency metrics.
Campaigns may opt into full transcript persistence with behavior.transcript_enabled;
the DAMAC campaign enables it and includes a transcription notice. Transcripts and local
recording references use the same protected row and campaign retention horizon. Raw
telephone numbers are never persisted. Do-not-contact and attempt tracking use a keyed
HMAC fingerprint. Structured call records and attempt history are assigned per-row
expirations using their campaign's data_retention_days; recording manifests use the
same campaign horizon, while suppression entries are retained. Run the
retention service beside the call worker so expired structured rows, recording WAVs, and
manifests are removed even while no calls are arriving:
PYTHONPATH=src .venv/bin/python -m speaking_agent.retention_worker \
--database data/speaking_agent.db \
--recording-directory data/recordings \
--interval-seconds 3600Populated databases created before per-row expiration fail closed. Migrate once with a reviewed horizon at least as long as every legacy campaign policy, then remove the option:
PYTHONPATH=src .venv/bin/python -m speaking_agent.retention_worker \
--database data/speaking_agent.db --legacy-retention-days 90 --onceAn opt-out is written through the suppression repository immediately upon recognition, before its spoken acknowledgement and call teardown.
Campaigns are runtime JSON files. Question order and wording come from configuration rather than a hardcoded sequence. The current deterministic grounding policy still contains property-domain evidence rules for fields such as intent, location, property type, price, timeline, and listing state. Add a pluggable or declarative domain policy before treating JSON alone as sufficient for a non-property campaign.
The default simulator, local voice, and LiveKit campaign is
campaigns/neoai_property_owner.json. That is where you
give the agent its overall goal and flexible scenario guidance:
The repository includes two examples:
campaigns/neoai_property_owner.json: DAMAC Lagoons seller qualification, market feedback, WhatsApp/document permission, and Yasir follow-up.campaigns/property_owner.json: legacy generic sell/rent property campaign.
{
"objective": "The measurable business goal and acceptable outcomes.",
"conversation_brief": "Who the agent is, who it is speaking with, useful context, and what success means.",
"conversation_guidelines": [
"How to sound and behave across every turn.",
"Answer what the person said before returning gently to the goal."
],
"scenario_playbook": [
{
"when": "A situation the person may introduce.",
"strategy": "What the agent should achieve or avoid, without prescribing exact words."
}
]
}objective, outcomes, fields, hard stops, and prohibited statements are firm policy.
conversation_brief, conversation_guidelines, and scenario_playbook guide Qwen's
judgment. They are explicitly passed as strategies rather than dialogue to quote, so the
agent can answer unexpected turns and vary its language while staying on goal. Exact
approved facts belong in faq_answers; standalone ASR/paraphrase routes to those facts
belong in faq_aliases; alternate follow-up wording belongs in question_variants.
FAQ aliases are exact punctuation-insensitive routes, so mixed FAQ-and-intent turns still
go through the model with the approved answer in context. Put a canonical FAQ question
in faq_answer_only when its approved answer should be spoken without immediately
appending the next qualification question, such as identity clarification or a complaint
about repetition. Compliant opening_variants are selected per session so repeated calls
do not always begin with identical wording.
The DAMAC campaign pursues selling intention, exact unit details, asking/minimum price,
grounded market feedback, WhatsApp permission, documents, and Yasir follow-up in that
order without treating the questions as a literal script. field_dependencies suppress
WhatsApp number and document questions when permission is declined. Market and outbound
workflow capabilities exist only when their approved endpoints are configured. An
explicit refusal ends the call without further probing, and an explicit future-contact
refusal additionally enters do-not-contact suppression.
During a session, the model receives bounded in-memory dialogue history for both the
owner and agent, plus the current stage, prior question counts, skipped fields, and known
structured facts. This lets it resolve references, remember what both sides said, and
avoid repeating an earlier answer or question. Configure the owner-turn bound with
behavior.conversation_memory_turns (default 12). This prompt window remains bounded.
The independent full dialogue is written to the retained call record only when
behavior.transcript_enabled is true.
To create another campaign:
- Copy
campaigns/property_owner.jsonand assign a uniquecampaign_id. - Configure the objective, introduction, full opening, and exact disclosure fragments.
- Define outcomes, outcome guidance, branch fields, field types, and one question per field.
- Configure hard stops, FAQs, prohibited statements, terminal/closing behavior,
voicemail text, attempt limits, retention, and model failure bounds.
Set
transcript_enabledexplicitly and include the reviewed notice/consent language required for the deployment. Setrecording_enabledexplicitly.trueenables dual-channel recording only when each controlled dispatch also carries a valid consent reference; otherwise the call is blocked before dialing. - Keep
controlled_test_modeenabled for controlled tests. Before disabling it, add a legally reviewed timezone and calling window. - Add regression scenarios and run the full suite.
Optional personalized_preamble configuration contains exactly three question
templates: recipient_confirmation with {recipient_name}, property_timing with
{property_reference}, and qualification without placeholders. The first template
must include all required AI/company disclosures. Metadata is opt-in; without it, normal
opening/opening_variants behavior is unchanged.
Campaign loading rejects undeclared outcomes, missing closings/questions/types, unsupported field types, invalid retention/attempt settings, and introductions missing required disclosures. It also rejects prohibited or human-identity claims in every directly spoken campaign surface: introductions, openings, variants, questions, closings, transfer/voicemail/error messages, and FAQ answers.
Choose the narrowest owning surface. Keep campaign-specific behavior in campaign JSON; change Python only when the behavior is reusable across campaigns or must be enforced outside the model.
| Need | Primary files | How to change it | Change it when |
|---|---|---|---|
| Change company identity, goal, wording, fields, FAQs, flow, voice style, or limits | campaigns/*.json |
Edit configuration, keep a unique campaign_id, then run campaign and conversation tests |
The behavior belongs to one campaign and fits the existing schema |
| Add or change campaign schema | src/speaking_agent/campaign.py, every campaign JSON, tests/test_campaign.py |
Add strict parsing/defaults and reject malformed or unsafe values | Multiple campaigns need a new declarative capability |
| Change DNC, transfer, outcome evidence, field grounding, or response safety | src/speaking_agent/policy.py, src/speaking_agent/text_safety.py |
Implement deterministic evidence rules and adversarial regressions | The rule affects compliance, persisted state, or terminal actions |
| Change adaptive turn progression or dialogue memory | src/speaking_agent/conversation.py, src/speaking_agent/domain.py |
Preserve policy ownership; update state only from validated evidence | The application must alter what objective/question comes next |
| Change Qwen prompting, context selection, parsing, or budget | src/speaking_agent/adapters/llm/qwen_mlx.py |
Keep sparse structured output, the 14,500-character cap, newest dialogue, and required policy context | Natural-language planning needs improvement without weakening policy |
| Add another LLM | src/speaking_agent/model.py, src/speaking_agent/adapters/llm/, composition entry points |
Implement ConversationModel; translate provider output to ModelInterpretation |
A platform cannot use MLX or a different quality/latency profile is needed |
| Add or tune ASR/TTS | src/speaking_agent/speech.py, src/speaking_agent/adapters/asr/, src/speaking_agent/adapters/tts/, src/speaking_agent/delivery.py |
Preserve application PCM/events and cancellation contracts | A new backend, language, voice, or measured quality target requires it |
| Change endpointing, short-speech handling, or local audio | src/speaking_agent/turn_detection.py, src/speaking_agent/local_voice_chat.py, src/speaking_agent/adapters/telephony/sounddevice_local.py |
Tune from captured consented scenarios; keep device/sample-rate behavior explicit | Audio evidence shows missed speech, false turns, echo, or latency problems |
| Change call lifecycle, barge-in, transfer, or cleanup | src/speaking_agent/voice_session.py, src/speaking_agent/transport.py |
Keep one CallSession owner and bounded cancellation/cleanup |
Behavior must be identical across local and LiveKit transports |
| Change LiveKit or controlled outbound SIP | src/speaking_agent/livekit_worker.py, src/speaking_agent/adapters/telephony/livekit_room.py, src/speaking_agent/call_cli.py, src/speaking_agent/outbound.py |
Keep SDK/SIP types inside adapters and preserve allowlist/suppression gates | Provider behavior or controlled-call requirements change |
| Change market-data integration | src/speaking_agent/market_data.py, composition entry points |
Preserve the normalized actual-transaction/current-listing distinction and approved HTTPS feeds | A reviewed provider or internal mapping service is available |
| Change lead classification/CRM delivery | src/speaking_agent/lead_workflow.py, src/speaking_agent/livekit_worker.py |
Keep raw numbers transient, use call_id idempotency, and preserve P1-P4/task semantics |
CRM or notification requirements change |
| Change persistence, privacy, retention, or metrics | src/speaking_agent/records.py, src/speaking_agent/recording.py, src/speaking_agent/suppression.py, src/speaking_agent/adapters/storage/sqlite.py, src/speaking_agent/metrics.py |
Migrate explicitly, keep raw numbers out, gate transcripts by campaign, and preserve atomic attempt checks | The structured record contract or reviewed retention policy changes |
| Change consented audio capture or recording retention | src/speaking_agent/audio_recording.py, src/speaking_agent/voice_session.py, transports, src/speaking_agent/retention_worker.py |
Preserve explicit per-call consent, private dual-channel artifacts, transport-confirmed playout, and expiry manifests | A reviewed quality/training workflow changes |
| Change operator workflows | src/speaking_agent/operator_cli.py, src/speaking_agent/retention_worker.py |
Add commands over repository interfaces rather than direct SQL | Operators need a repeatable inspection or maintenance action |
- Start with a failing scenario or adapter-contract test in the matching
tests/test_*.py. - Prefer a campaign-only edit when the schema can express the requirement.
- Put critical decisions in policy/application code; treat model output as untrusted suggestions.
- Add a protocol or abstraction only when a second implementation or real duplication requires it.
- Run focused tests after the first edit, then the complete quality gate above.
- For audio, model, LiveKit, SIP, or platform work, also run the native integration on the target hardware. Wheel installation or mocked tests do not establish support.
- Update this README for user-visible behavior and update
ROADMAP.mdwhen an item is completed, split, reprioritized, or newly discovered.
Minimum validation by change type:
| Change type | Required evidence |
|---|---|
| Campaign content/schema | Campaign tests plus relevant conversation scenarios |
| Policy/state/lifecycle | Focused regression, full suite, and failure-path coverage |
| LLM/ASR/TTS adapter | Contract tests plus one real model smoke on supported hardware |
| Local audio/turn detection | Unit tests plus microphone/speaker round trip on the target device |
| LiveKit/SIP | Adapter tests plus controlled room/call evidence; never use an unapproved number |
| Storage/privacy | Migration, concurrency, retention, and no-sensitive-data assertions |
| New platform/profile | Clean native install, full gate, hardware round trips, and published measurements |
- LLM: implement
ConversationModel.prepare,interpret, andclose; translate output toModelInterpretation. - ASR: implement
SpeechRecognizer; emit cumulative partial/finalTranscriptEventobjects. - TTS: implement
SpeechSynthesizer; emit signed 16-bit little-endian PCMAudioFrameobjects and support cancellation. - Telephony: implement
CallTransport, including queue clearing and playout draining. - Storage: implement
CallRepository, including suppression, attempt, and retention operations.
Mocks cover every boundary. They are the supported path for testing without models, audio hardware, LiveKit, or telephony.
- LiveKit room audio is verified locally in both directions. PSTN integration is implemented but not exercised because no project credentials, SIP trunk, or controlled number were supplied.
- Qwen3-ASR currently receives a completed detected utterance, then streams decoder events. It does not incrementally encode live PCM as it arrives.
- The included energy turn detector is intentionally simple. Real telephone noise should be measured before selecting or tuning a stronger VAD/noise-cancellation adapter.
- Answering-machine detection uses explicit machine phrases and otherwise defaults to human to avoid discarding long human responses. Production AMD needs measured data.
- MLX cancellation is cooperative. Teardown waits for a bounded grace period and quarantines a stalled adapter; a native Python worker thread may finish later.
- Human-identity protection combines campaign-load validation and runtime regex-based filtering. Known variants are covered, but new paraphrases should be added as adversarial tests; truthful AI/automation disclosure must never be weakened.
- The Qwen prompt is capped at 14,500 characters by dropping oldest dialogue and then optional guidance. Required policy/state/current-turn context is never silently dropped; an irreducible oversized campaign fails explicitly and should be shortened.
- Campaign wording and question flow are configurable, but deterministic extraction and
outcome grounding currently include property-domain logic in
policy.py. Introduce a tested domain-policy boundary before deploying a materially different campaign type. - Greedy Qwen TTS removes sampling variance and repeated-text output is deterministic on
the verified checkpoint. Perceptual consistency across different sentences still
requires listening benchmarks; use
--tts-temperature 0for the most stable profile. - Consented audio recording stores sensitive raw PCM and a minimal manifest. Application- level encryption, de-identification, annotation, approval workflow, and training-data export are not implemented; owner-only permissions and retention are not substitutes for those production controls.
- SoundDevice reports callback-consumed agent PCM precisely. LiveKit currently reports completed playout batches; when a LiveKit turn is interrupted, its agent-channel prefix may be omitted because the SDK path does not expose an exact played-sample cursor.
- The example company, wording, retention, and policy values are demonstrations, not legal advice. Production use requires authoritative review for identity, consent, recording, calling times, retention, transfer, and suppression rules.
- No queue, broker, microservice split, retry farm, or multi-machine deployment is added; the current requirement is one process and one controlled call.
- DLD/Property Monitor/listing credentials and exact identifier mappings were not supplied, so live market accuracy and external webhook delivery remain unverified.