RLM_Wikidata is a Recursive Language Model (RLM) harness over a frozen Wikidata graph. The model answers a question by writing Python in a REPL: it searches entities, follows relations, reads qualifiers and references, keeps intermediate results in variables, and returns an answer checked exactly against the graph.
Wikidata describes more than 100 million entities in over 300 languages, and AI systems have no good way to explore it. SPARQL takes expertise and pasting graph data into a context window degrades as the question grows. Recursive Language Models (RLMs; Zhang, Kraska & Khattab, 2026) fit this task better: the model explores the graph in code and calls itself on the parts that matter.
This harness generated Wikidata-Search-Traces, an open corpus of 10,235 reasoning trajectories over Wikidata, available on Hugging Face.
- One cell per turn. The model replies with one Python block. The harness runs it in a persistent namespace, so variables carry over between turns, and returns the printed output, truncated to 10,000 characters. The model ends the run with
FINAL(value). - Graph functions. The model reads the graph only through 13 functions:
search_entity,search_property,claims,references,describe,labels,label,descriptions,edges,count_edges,degree,nameandhas. Every call is logged. - Sub-calls.
llm_queryandllm_mapsend text to a second instance of the model, for semantic judgements over material the run has already read. Sub-calls have no graph access. - Grounding checks. The harness rejects a final answer whose shape does not match the answer format, or that names an entity the run never read.
- Full archive. Every run keeps the messages, reasoning, executed code, graph reads, sub-calls, timing and the final answer.
- Exact scoring. Every answer is an entity (QID), a year, a quantity, a number, a string or a list of entities, so a deterministic scorer compares it with the reference answer. No language model judges correctness.
We worked over a frozen Wikidata graph. The harness works with any OpenAI-compatible endpoint; the reference configuration serves Qwen3.8-27B-FP8 with vLLM on one GPU.
We evaluated four models on 100 questions built with the same pipeline as the corpus: 50 single-entity and 50 multi-hop, in tasks/eval/, scored by exact match with no judge. Each model answers either through this harness or as a tool-calling agent that calls the same 13 graph functions directly. With the model held fixed, the harness improves both models we ran under both interfaces: gpt-6-luna from 49 to 61 correct answers and Qwen3.8-27B from 60 to 74. The technical report gives the full protocol and analysis.
| System | Model | Setup |
|---|---|---|
| Qwen3.8-27B + RLM | Qwen3.8-27B (FP8, served with vLLM on one H100) | RLM harness |
| gpt-6-luna + RLM | gpt-6-luna | RLM harness |
| Qwen3.8-27B agent | Qwen3.8-27B (FP8, served with vLLM on one H100) | tool-calling agent |
| glm-5.3-flash agent | glm-5.3-flash | tool-calling agent |
| gpt-6-luna agent | gpt-6-luna | tool-calling agent |
| gemini-3.1-flash-lite agent | gemini-3.1-flash-lite | tool-calling agent |
- Shared: every system uses the same 13 graph functions over the same frozen graph, with low reasoning effort. Every system gets the same instruction: "The question has an answer. Do not stop until you find one."
- RLM harness: the model writes Python in a REPL and keeps intermediate results in variables. It can call itself on sub-problems. Temperature 1, up to 100 turns and 80 sub-calls maximum.
- Tool-calling agents: openai-agents with LiteLLM. They call the graph functions directly, with tool output truncated at 10,000 characters, up to 100 calls and $0.50 per question maximum.
| Model | Interface | Single-entity | Multi-hop | Total | Cost / question | Tokens / question (median) | Seconds / question (median) |
|---|---|---|---|---|---|---|---|
| Qwen3.8-27B | RLM harness | 46/50 | 28/50 | 74/100 | $0.04 ¹ | 35,754 | 25 ² |
| gpt-6-luna | RLM harness | 42/50 | 19/50 | 61/100 | $0.013 | 24,942 | 18 |
| Qwen3.8-27B | tool-calling agent | 40/50 | 20/50 | 60/100 | $0.025 ¹ | 32,072 | 17 ² |
| glm-5.3-flash | tool-calling agent | 37/50 | 13/50 | 50/100 | $0.067 | 30,478 | 33 |
| gpt-6-luna | tool-calling agent | 40/50 | 9/50 | 49/100 | $0.005 | 24,924 | 16 |
| gemini-3.1-flash-lite | tool-calling agent | 31/50 | 10/50 | 41/100 | $0.19 | 460,426 | 104 |
For the API models, cost is input tokens × input price + output tokens × output price, with prices from LiteLLM's model price table; reasoning tokens count as output tokens.
¹ Estimated cost of batched serving on one H100 in Google Colab. ² Qwen3.8-27B served with vLLM on one H100, processing one question at a time.
| Path | Contents |
|---|---|
scripts/rlm_loop.py |
The harness: runs one question and archives the trajectory |
scripts/wd_graph_env.py |
The graph environment: the functions the model calls |
scripts/bench/ |
Batch runner, scorer and summary |
tasks/eval/ |
The 100 evaluation questions with their reference answers |
config/runs/reference.env |
The settings used for the evaluation |
docs/running.md |
How to set up and run the harness |
docs/graph-format.md |
The graph files the environment reads |
The harness needs Python 3.13, a Wikidata graph in the format described in docs/graph-format.md, and any OpenAI-compatible endpoint: a local server such as vLLM, or a hosted API. docs/running.md covers installation, serving a model, the settings, and how to run and score the evaluation.
uv sync
set -a; source config/runs/reference.env; set +a
export WIKIDATA_GRAPH_DIR=/path/to/graph
export BASE_URL=http://127.0.0.1:8000/v1 API_KEY=EMPTY MODEL_NAME=<served model name>
.venv/bin/python scripts/bench/run_batch.py --families eval --workers 1 --tag eval
.venv/bin/python scripts/bench/summary.py results/runs/<batch>