Skip to content
cosmol-studioPublic

About

Pure-Rust cheminformatics toolkit and RDKit-compatible alternative for SMILES/SMARTS, SDF, fingerprints, substructure search, ETKDG, UFF/MMFF, and Python workflows.

Topics

Resources

Stars

30 stars

Watchers

1 watching

Forks

Latest commit

Β 

History

179 Commits

Folders and files

Repository files navigation

COSMolKit β€” Rust-native cheminformatics toolkit

coverage workflow badge codecov badge crates.io badge docs.rs badge pypi badge npm badge MIT license badge

COSMolKit is a Rust-native cheminformatics and structural biology toolkit with first-class Python bindings. It provides molecular graph operations, SMILES/SMARTS and molecular file workflows, 2D depiction, native 3D conformer generation, UFF/MMFF optimization, fingerprints, molecular descriptors, InChI, batch processing, and protein structure APIs. Selected workflows are also available directly in the browser through the COSMolKit Web Tools.

For supported cheminformatics operations, RDKit-compatible behavior is treated as the correctness floor. COSMolKit uses boundary-scoped parity claims: a feature is considered parity-covered only when its documented reference surface passes the required exact or numerical comparisons. Fixed reference oracles, source-backed implementations where reference semantics require them, committed regression corpora, and explicit capability boundaries are used together; aggregate success rates or approximate similarity are not treated as substitutes for behavioral parity.

COSMolKit combines a native Rust API with Python interfaces designed for array-oriented scientific and machine-learning workflows. Molecular graphs, coordinates, fingerprints, bounds matrices, and structural data are exposed in forms suitable for NumPy, PyTorch, dataset processing, and model-building pipelines.

COSMolKit 0.5.0

  • Modular Rust architecture: separate crates own model values and domain algorithms; cosmolkit remains the public entry point and molecule runtime.
  • Cargo features: select functionality such as core, bio, fingerprints, or reaction; prerequisites are included. Default is full; see the Rust feature guide.
  • WebAssembly support: run the same Rust chemistry engine in the browser through JavaScript bindings, without a separate algorithm implementation.
  • One API, three languages: shared operations, defaults, results, and error categories, with idiomatic names for Rust, Python, and JavaScript. See Public API Design for cross-language naming rules and shared API contracts.

Each public API is one logical entry with three projections:

Logical name Rust Python JavaScript
from_smiles Molecule::from_smiles Molecule.from_smiles Molecule.fromSmiles
molecular_weight mol.molecular_weight() mol.molecular_weight() mol.molecularWeight()
to_smiles mol.to_smiles() mol.to_smiles() mol.toSmiles()
with_hydrogens mol.with_hydrogens() mol.with_hydrogens() mol.withHydrogens()
tpsa mol.tpsa() mol.tpsa() mol.tpsa()

Learn once, then move between Python notebooks, native Rust applications, and browser-based JavaScript without relearning the chemistry API.

Documentation

Validation Status

COSMolKit treats parity as source-backed semantic equivalence within explicitly documented boundaries, not as statistical agreement of final outputs. Compatibility-critical chemistry is implemented as a line-by-line, source-backed port with explicit operation contracts and traceable correspondence to pinned upstream code. Validation corpora verify that port; they are not used to iteratively tune heuristic reimplementations until outputs happen to agree.

The comparison boundary therefore extends well beyond final strings. Covered surfaces compare exact bytes, bits, return status, complete atom and bond state, stereochemistry, derived state and invariants, RNG state, seed handling, and random draw sequences where stochastic behavior is part of the contract, every matrix entry, coordinates, energies, and every gradient component where applicable. Discrete results must match exactly; declared numerical tolerances reach 1e-8 for matrix entries and 1e-6 for coordinates, energies, and gradients. 99% or 99.9% agreement remains unfinished when any covered mismatch exists.

0.5.0 validation is pending. The published ChEMBL 37, 5,000-record and 152-record results in VALIDATION.md come from 0.3.0. They are retained as historical evidence, not a claim that 0.5.0 has passed the same comparisons. Current validation must record the tested implementation, reference versions, options, counts and actual outcomes.

Installation

pip install cosmolkit

Core Concepts

  • Value-style molecules: methods such as with_hydrogens(), without_hydrogens(), with_kekulized_bonds(), and with_2d_coordinates() return new molecule values, keeping topology-changing operations explicit and preventing derived chemistry state from being silently invalidated.
  • Explicit mutation: in-place Molecule operations always end with _. The trailing underscore has no other public Molecule meaning.
  • Explicit errors: invalid input and unsupported behavior are surfaced as errors instead of silent fallbacks.
  • Batch-native processing: MoleculeBatch keeps input order, supports structured per-record failures, and can run batch transforms and exports with configurable parallelism.
  • Array-friendly data access: coordinates, bounds matrices, fingerprints, and graph features are exposed in forms that fit Python numerical workflows.
  • Source-backed 3D workflows: conformer generation and UFF/MMFF optimization are available through the public Python API, and atom chiral tags can be assigned from a selected 3D conformer with pinned-RDKit parity.

Value-Style Transformations

Normal molecule operations return new objects and do not mutate their inputs. This follows the same explicit-dataflow direction as modern dataframe libraries: users can reason about each transformation as a new value while COSMolKit can share unchanged internal storage efficiently.

import cosmolkit as ck

mol = ck.Molecule.from_smiles("CCO")
mol_h = mol.with_hydrogens()

assert mol is not mol_h

Python Quick Start

import cosmolkit as ck

mol = ck.Molecule.from_smiles("c1ccccc1O")
mol_2d = mol.with_2d_coordinates()

print(mol_2d.to_smiles())
print(mol_2d.coordinates_2d())

mol_3d = mol.with_hydrogens().with_3d_conformer()
print(mol_3d.coordinates_3d())

svg = mol_2d.to_svg(width=400, height=300)
mol_2d.write_png("phenol.png", width=400, height=300)

fp = mol.fingerprint_morgan()
print(fp.on_bits())

atom_pair = mol.fingerprint_atom_pair()
print(atom_pair.on_bits())

layered = mol.fingerprint_layered()
print(layered.on_bits())

pattern = mol.fingerprint_pattern()
print(pattern.on_bits())

stereoisomers = list(ck.Molecule.from_smiles("FC(Cl)Br").enumerate_stereoisomers())
print([isomer.to_smiles() for isomer in stereoisomers])

batch = (
    ck.MoleculeBatch.from_smiles_list(
        ["CCO", "c1ccccc1", "CC(=O)O"],
    )
    .with_parallel_jobs(8)
    .with_progress_bar(False)
)

prepared = batch.with_hydrogens().with_2d_coordinates()
print(prepared.valid_mask())
print(prepared.to_smiles_list())

prepared.write_images("molecule_images")

Protein Structures

Use BioStructure for complete PDB/mmCIF structural data, including modeled proteins, nucleic acids, ligands, waters, entities, models, and metadata. Use Protein only when an amino-acid-only projection is intended.

import cosmolkit as ck
from pathlib import Path

structure = ck.BioStructure.from_pdb(Path("complex.pdb").read_text())
print(structure.num_models(), structure.num_chains(), structure.num_atoms())

# Structural format conversion remains on the complete structural value.
mmcif_text = structure.to_mmcif()
roundtrip = ck.BioStructure.from_mmcif(mmcif_text)
print(roundtrip.num_atoms())

The protein projection is explicit and leaves the full structure available:

protein = structure.protein()

print(protein.num_chains())
print(protein.num_residues())
print(protein.num_atoms())

SDF and Dataset Workflows

SdfDataset builds a lightweight index of SDF record byte ranges, so individual records and chunks can be read without loading an entire file into memory. Molfile-only readers such as Molecule.read_mol() follow RDKit MolFromMolBlock boundaries: they stop after the first M END line and leave trailing SDF data fields to the SDF APIs.

import cosmolkit as ck

dataset = ck.SdfDataset.open("library.sdf")
print(len(dataset))

record = dataset[0]
mol = record.molecule()

for batch in dataset.batches(size=1024, errors="keep", n_jobs=8):
    smiles = batch.to_smiles_list()

Conformer Generation And Optimization

import cosmolkit as ck

mol = ck.Molecule.from_smiles("CC(=O)NC").with_hydrogens()

params = ck.EmbedParams(random_seed=0xF00D, num_threads=1)

embedded = mol.with_3d_conformer(params)
print(len(embedded.conformers_3d()))
print(embedded.coordinates_3d())

multi = mol.with_3d_conformers(5, params)
print(len(multi.conformers_3d()))

uff = embedded.uff_force_field()
print(uff.energy())
outcome = uff.minimize_(max_iterations=200)
print(uff.energy())

mmff = embedded.mmff_force_field()
print(mmff.energy())
outcome = mmff.minimize_(max_iterations=200)
print(mmff.energy())

with_3d_conformer() uses source-backed distance-geometry embedding for trusted molecular graphs: molecules without explicit hydrogens are embedded as heavy-atom-only conformers instead of failing or automatically adding hydrogens. Calling with_hydrogens() first is recommended for all-atom geometry, force-field optimization, and hydrogen-bond-sensitive workflows. Coordinate-only inputs such as XYZ blocks do not contain a bond topology and are not valid ETKDG inputs until a trusted graph has been constructed.

Feature Areas

  • Molecular graph construction and inspection
  • SMILES parsing and writing
  • MOL/SDF reading and writing
  • MOL2 reading with RDKit-style Mol2ParserParams
  • XYZ block reading
  • Four scalar InChI APIs with exact source-defined official-C/RDKit parity and structured errors
  • Source-backed 3D atom-chiral-tag assignment
  • Typed potential-stereo analysis and lazy, source-ordered stereoisomer enumeration with exhaustive, bounded random, uniqueness, enhanced-group, and optional embedding controls
  • Hydrogen transforms and Kekulization
  • Sanitization and chemistry problem detection
  • Source-backed Rust tautomer enumeration, canonical selection, scoring, callbacks, current and V1 transform catalogs, and ordered result provenance
  • 2D coordinate generation and SVG/PNG depiction
  • Native 3D conformer generation with DG/KDG/ETDG/ETKDG parameter presets
  • Read-only molecular alignment/RMSD measurement and explicit value-style or trailing-underscore coordinate alignment
  • UFF/MMFF optimization of generated or imported 3D conformers
  • Morgan, MACCS, RDKit topological, Avalon, Pattern, AtomPair, and Topological Torsion fingerprints for the validated exact-parity branches, including sparse/count forms, provenance, tautomer-aware Pattern hashing, 2D/3D AtomPair distances, and ordered batch execution
  • Source-backed RDKit Layered fingerprint 0.7.0 with all six active layers, roots, masks, seeded atom counts, and ordered batch execution; this family retains upstream's experimental classification
  • Distance-geometry bounds matrices
  • Substructure matching and SMARTS parse metadata
  • Ordered batch transforms and exports
  • Python pickle round-tripping for Molecule
  • Complete PDB/mmCIF BioStructure parsing, Gemmi-aligned mmCIF writing, protein projections, and explicit structure-to-molecule conversion
  • Support-status metadata for public features

Design Principles

COSMolKit aims to be Python-friendly, batch-friendly, and suitable for model-building workflows.

  • Correctness comes before breadth.
  • Public transforms use value semantics.
  • Mutation-capable workflows are explicit.
  • Fail-closed capability boundaries: a separately named capability outside documented support returns a structured error rather than fabricated chemistry. This is an API design rule, not an accepted mismatch within a supported feature.
  • RDKit-parity behavior is the correctness floor for supported cheminformatics features.
  • High-throughput APIs should preserve input order and expose per-record failures.
  • Reference semantics come before heuristic approximation; semantic debt is treated as a correctness risk.

Examples

Python examples live in python/examples/. For the current InChI interface, see python/examples/inchi_roundtrip.py.

Development

Small focused Rust test filters may use the default debug profile while iterating:

cargo test -p cosmolkit-core <test-filter>

Large local runs, parity suites, and CI tests should use release mode with the runtime strict checks enabled on cosmolkit:

cargo test -p cosmolkit-core --release
cargo test -p cosmolkit --release --features op-contracts-strict

Detached core algorithms have no runtime-check features. Release-mode testing keeps operation contracts and runtime invariants enabled on cosmolkit through op-contracts-strict; optimized release builds for distribution use default features unless explicit runtime checks are requested.

Roadmap

Status labels:

  • βœ… stable public functionality within its documented supported scope
  • πŸ§ͺ public experimental feature; available, but its behavior or API may change
  • 🚧 planned or not yet public

The βœ… status applies to the documented COSMolKit scope. It does not claim that every API or input branch from an upstream reference library is implemented; separately named upstream capabilities outside that scope are not represented as implemented. Every path inside a parity-covered boundary is still required to match; individual failing rows cannot be reclassified as out of scope.

Chemistry Core

Goal: keep the supported molecular core correct before expanding breadth.

  • βœ… Molecule, atom, and bond graph model
  • βœ… SMILES parsing
  • βœ… SMILES writing with RDKit-style writer options for supported branches
  • βœ… Ring perception, valence handling, aromaticity, and Kekulization
  • βœ… Hydrogen addition and removal
  • βœ… Sanitization for supported chemistry workflows
  • βœ… Tautomer enumeration, canonical selection, scoring, and source-defined stereochemistry and isotopic-hydrogen options
  • βœ… Stereochemistry inspection for supported atom and bond states
  • βœ… Typed potential-stereo analysis and lazy stereoisomer enumeration through the pinned RDKit Python behavior, including arbitrary-width counts, seeded random generation, enhanced stereo groups, uniqueness, and optional embedding
  • βœ… Atom chiral-tag assignment from selected 3D conformers, with exact pinned-RDKit assignChiralTypesFrom3D parity across 77 fixed full-state oracle records
  • βœ… Distance-geometry bounds matrices
  • βœ… Native 3D conformer generation and UFF/MMFF post-optimization for supported molecules
  • βœ… Molecule.to_inchi(), Molecule.to_inchi_key(), inchi_to_key(), and Molecule.from_inchi() for source-defined behavior; official-C undefined allocation behavior returns a structured error
  • βœ… Source-backed Morgan, MACCS, RDKFingerprint/topological, Avalon, Pattern, AtomPair, and Topological Torsion fingerprints for their validated exact-parity branches, including typed provenance, count/bit forms, tautomer-aware Pattern hashing, 2D/3D AtomPair distances, Tanimoto similarity, and ordered batch execution
  • πŸ§ͺ Source-backed RDKit Layered fingerprint 0.7.0, including all six active layers, roots, masks, seeded counts, and ordered batch execution; upstream classifies this fingerprint family as experimental
  • βœ… Substructure matching and Python SMARTS parse metadata
  • βœ… Molecular descriptors: weight/formula, H-bond and Lipinski counts, Crippen/TPSA/QED, connectivity Chi, Hall-Kier/Kappa/Phi, ring and stereo counts, MQN, Labute ASA, and SlogP/SMR VSA for the documented parameter space

File I/O and Depiction

Goal: make common molecule import, export, and visualization workflows usable from Python.

  • βœ… MOL/SDF reading
  • βœ… MOL2 reading
  • βœ… XYZ block reading
  • βœ… SDF dataset indexing for large files
  • βœ… SDF writing for supported V2000/V3000 branches
  • βœ… PDB block to molecule conversion
  • βœ… mmCIF block to molecule conversion through the same molecule-conversion profile
  • βœ… 2D coordinate generation
  • βœ… SVG drawing
  • βœ… PNG export
  • βœ… RDKit-style visual parity testing for supported depiction output
  • 🚧 Annotation overlays and richer drawing customization
  • βœ… 3D conformer generation and embedding APIs

Batch-Native Workflows

Goal: make high-throughput molecule preparation and export a core product identity.

  • βœ… Ordered MoleculeBatch.from_smiles_list()
  • βœ… Batch transforms for sanitization, hydrogens, Kekulization, and 2D coordinates
  • βœ… Configurable parallelism with with_parallel_jobs()
  • βœ… Configurable progress display with with_progress_bar()
  • βœ… Per-record errors, valid masks, and error reports
  • βœ… Batch SMILES, image, and SDF export paths
  • βœ… Golden parity tests for parallel batch behavior
  • 🚧 More streaming and chunked dataset workflows

Protein and Structural Biology

Goal: provide practical Biopython-like structure workflows without forcing users through low-level structural tables.

  • βœ… Protein.from_pdb() / Protein.from_mmcif() high-level entry points
  • βœ… BioStructure.from_pdb() / BioStructure.from_mmcif() complete-structure entry points and mixed-structure hierarchy traversal
  • βœ… Protein chain, residue, and atom iteration
  • βœ… Protein-only projection from broader structural data
  • βœ… PDB/mmCIF structural parsing
  • βœ… Gemmi-aligned BioStructure mmCIF serialization and file writing
  • 🚧 Selection utilities for chains, residues, atoms, and neighborhoods
  • 🚧 Ligand, nucleic-acid, and mixed-structure ergonomic APIs

Python API and ML Readiness

Goal: expose verified molecular behavior through a practical Python interface.

  • βœ… Stable value-style mutation contract for public molecule transformations
  • βœ… Graph, coordinate, fingerprint, descriptor, and bounds-matrix accessors
  • βœ… Python examples for drawing, SDF-to-SMILES, pickle round-tripping, batch processing, and proteins
  • βœ… Type stubs and documentation coverage
  • 🚧 Stable model-ready graph exports
  • 🚧 NumPy / PyTorch oriented adapters
  • 🚧 Molecular tokenization and AI-native geometry helpers

Browser and Deployment

Goal: make validated COSMolKit functionality usable without requiring a local Python or Rust installation.

  • βœ… COSMolKit Web Tools for browser-based molecular workflows
  • βœ… Browser-native deployment of selected COSMolKit functionality through WebAssembly
  • 🚧 Broader JavaScript bindings
  • 🚧 Expansion of browser-native chemistry and structural-biology workflows

Acknowledgments

COSMolKit includes Rust ports of code and algorithms from:

License

COSMolKit is licensed under the MIT License. Vendored sources and externally derived test fixtures retain their upstream copyright and license terms as documented beside those files.

About

Pure-Rust cheminformatics toolkit and RDKit-compatible alternative for SMILES/SMARTS, SDF, fingerprints, substructure search, ETKDG, UFF/MMFF, and Python workflows.

Topics

Resources

Stars

30 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages