Standardizing nutrition science for preventive medicine.
Documentation · Cookbook · Validation report
Engineering Considerations (we found this quote that supports what we are trying to do from an total expert in the field.)
Given the functions of the gastrointestinal system discussed earlier, we turn now to a consideration of anatomic features needed to support these functions. In this discussion, the gastrointestinal system can be thought of as a machine (Figure 1–1) in which distinct portions conduct the various processes needed for assimilation of a meal without uptake of significant quantities of harmful substances or microorganisms.
— Barrett, K. E. (2014). Gastrointestinal physiology (2nd ed.). McGraw-Hill Education, p. 3.
biology-as-code models what happens to a meal — digestion, absorption, and the
metabolic pathways it drives — as inspectable, versioned, provenance-tracked code.
meal → digestion → pathway claims, fail-closed
Nutrition data loses its origins as it travels. A number measured once in one lab becomes a label value, becomes a database entry, becomes an input to a score, and by the end nobody can say what it was or how well it was evidenced. This package takes the opposite stance: every value carries where it came from, missing data stays missing instead of quietly becoming zero, and the rules are queryable data rather than assumptions buried in code.
It is a zero-dependency Python package (3.11+) and it works on its own today.
Not medical advice. Teaching and research software — not a clinical decision-support system.
On the name. Code Biology (Barbieri and others) is an existing field that studies organic codes in living systems. This project is unrelated to it. The name here marks a methodological stance rather than a claim about semiotics: that nutrition and pathway models should be written like software — versioned, tested, provenance-tracked, fail-closed. That field is descriptive literature; this is a prescriptive tool. See docs/naming.md.
Three working papers on federal nutrition data infrastructure are now published and citable on Zenodo — announcement thread:
These are policy and infrastructure companions to this repository: the papers argue the case for provenance-tracked nutrition data, and this package is a working implementation of it.
pip install biology-as-codeOr from source: git clone this repo and pip install -e ".[dev]".
from biology_as_code import simulate_meal, list_pathways, pathway_activities, fed
result = simulate_meal(carbs_g=55, protein_g=35, fats_g=18, fiber_g=20)
print(result.absorbed_macros_g)
# {'carbs': 51.894, 'protein': 33.041, 'fats': 15.102}
print(list_pathways()[:5])
# ['glycolysis', 'tca_cycle', 'etc_oxphos', 'beta_oxidation', 'gluconeogenesis']
print(pathway_activities(fed())['glycolysis'])
# 0.868Two runnable examples ship with the repo:
python examples/python/run_meal.py
python examples/python/fed_vs_fasted.pyLogging is quiet by default; export BIOLOGY_AS_CODE_LOG=DEBUG turns on digestion traces.
- 40 metabolic pathway graphs — glycolysis, TCA, β-oxidation, ketogenesis and
ketolysis, AMPK·mTORC1·SREBP nutrient sensing, and more. Fetch with
get_pathway(...), render withvisualization.pathway_to_mermaid. - 10 digestion machines — the GI tract from oral to colon as versioned, inspectable
state graphs:
list_machines(),trace(),run_digestion(...). - 47 LAW-SPEC law cards — the constitution as queryable data.
law_card("LAW-004")returns its System, Organ, Gate, Bound, Conditions, and relation. - Meal simulation —
simulate_meal(...)plus fed, fasted, and exercise scenarios. - Bundled data — meal fixtures, a vitamins registry, personas, and iron/colon/law data.
- Provenance throughout —
all_sources()andpubmed_url()surface every citation. No fabricated data, and no network calls by default.
Not included: the patent-pending product/meal score (an open hook only), and the book text.
- Label ≠ dose — printed milligrams are not delivered dose.
- Gate ≠ bound — whether something can happen is a separate question from how much.
- Four seats — host, partner, stage, clock, in one law envelope.
- L1→L5 — matrix → nutrient → mechanism → physiology → outcome, with no tunnels between levels.
- Empty beats fake — missing data is
UNEVALUABLE, never a green score.
The full reasoning is in the constitution.
This package is a reference implementation of FDP-1: Food Data Provenance Declaration — a minimal, RFC-style specification for declaring where a nutrient value came from and how well a score built on it is validated. FDP-1 wraps any existing system (Nutri-Score, Health Star Rating, Nutri-Grade, Food Compass, or a proprietary score) without modifying it.
- Seven fields on a value, five on a score, one rule — the weakest-link rule (§3.1): a score's provenance grade equals its lowest-graded input.
OPENvsNONE(§4) — not known versus known-absent or not-applicable. These are different claims and the spec keeps them apart.nutrient_refresolves to the canonical CDNO food-composition vocabulary, with ChEBI, FDC, and INFOODS accepted as alternate keys. It does not resolve toMASTER_CROSSWALK.tsv, the nutrient→metabolite join this repo hosts as a separate downstream layer.
The specification, its reference validator, and the worked example live in the canonical
fdp-1 repository, published as a citable Standard
on Zenodo: doi:10.5281/zenodo.21613721
(concept DOI — always resolves to the latest version).
Small teaching packets, each built to make one mechanism visible:
| File | Teaching point |
|---|---|
spinach_salad_zero_fat.json |
Fat-vehicle gate closed |
spinach_salad_with_oil.json |
Same cargo, plus a lipid partner |
lentils_with_tea.json |
Iron bound narrowed by tannin |
lentils_with_ascorbate.json |
Iron bound expanded by ascorbate |
almond_whole.json / almond_flour.json |
Matrix intact versus destroyed |
Files marked status: stub are placeholders awaiting real cargo and partner data — oats,
breads, rice, juice, salmon, dairy, tofu, UPF snacks, supplements, and more. See
examples/foods/ for the full set. To add one, copy
_template.json, keep the schema, and leave fields
"open" until you have real values. That last part is the whole point.
docs/ # public documentation site
schemas/ # packet + claim + relation subset
examples/
foods/ # teaching food packets (gates / claims)
claims/ # claim audit fixtures
units/ # teaching UNIT fixtures
python/ # runnable scripts
src/biology_as_code/
data/fixtures/meals/ # full meal JSON, ships in the wheel
data/fixtures/ # vitamins, personas
pathways/ dig/ simulation/ visualization/
Meals live in exactly one place (src/biology_as_code/data/fixtures/meals/). Foods are a
separate concept in examples/foods/ — teaching packets, not meals.
- Package architecture
- Add a pathway — template, checklist, and integration check
- GEM primer — genome-scale metabolic models, and how this differs
- A note on the name — why this is unrelated to Barbieri's Code Biology
- Related projects — what this is, what it is not, and what we consume
- Contributing · Roadmap · Changelog
Seventeen neighbouring projects — ontologies, metabolic modelling stacks, food-composition initiatives, and health schemas — are assessed in docs/related-projects.md, each with a verified licence and a stated position: upstream (we consume it), adjacent (cite, don't vendor), watch, or unrelated. The short version:
| Layer | Position |
|---|---|
| FoodOn, CDNO, FOBI, CompTox | Upstream. Already named by the claim vocabulary — we resolve to them, we do not rebuild them. |
| COBRApy, Tellurium, libSBML, PySB | Adjacent. Different altitude, and COBRApy is GPL-2.0 — a dependency this Apache-2.0 package cannot take. |
| PTFI | Watch. Molecular characterisation is a layer below claims; the open question is whose identifiers everyone joins on. |
| Bioregistry | Upstream — and it has no prefix for FoodData Central or INFOODS, two namespaces FDP-1 already accepts. |
That page also records two verified findings worth knowing outside this repo: FoodOn ships
no FoodData Central mapping (its USDA cross-references are to PLANTS, a taxonomy), and
CDNO's licence is CC BY 3.0 in the ontology header while its GitHub LICENSE file says
CC0. The licence of an artefact is what the artefact declares.
Closest sibling: this repo models what a body does with a meal, while Awesome Internet of the Body is the companion list of sensors and apps that measure the body itself — CGMs, wearables, FHIR, Open Humans, and others.
Biology as Code — Standardizing Nutrition Science for Preventive Medicine is in progress and not yet published. There is nothing to buy or preorder yet, and the manuscript is not in this repository; it will be released separately as a commercial book.
This repo is its open companion and is fully usable on its own. Issues are welcome for schemas and examples — not for the book text.
Alpha. The Python package works today and installs from source; the schemas and food objects will keep growing. Nothing here depends on the book being finished.
- Code: Apache-2.0, with a patent non-assertion covenant in PATENTS.md.
- Schemas and examples: see LICENSE-SAMPLES.md — permissive reuse with attribution.
- Book text, figures, and brand: © the author, all rights reserved unless separately licensed.
If you use biology-as-code in your work, please cite the archived release:
Murff, P. (2026). Biology as Code: an open, provenance-tracked toolkit for meal digestion and metabolic-pathway modeling (v0.1.0) [Software]. Zenodo. https://doi.org/10.5281/zenodo.21536449
A CITATION.cff is included, so GitHub's Cite this repository button
works as well.
The book — Biology as Code: Standardizing Nutrition Science for Preventive Medicine — is a separate, related work, coming soon: the book page has the cover and the pitch.
Cover art is a draft. The glucose chemistry shown on it is not yet corrected.