Skip to content

Convert: generic HDF5 matrix and delimited text files #25

Description

@Claptar

Part of #20. Depends on #21. Covers the remaining scanpy/anndata readers: read_hdf, read_csv, read_text.

Generic HDF5 (read_hdf)

scanpy's read_hdf(file, key) treats one HDF5 dataset as a dense X.

adata convert m.h5 -o m.h5ad --key /path/to/matrix \
    [--obs-names-key /path] [--var-names-key /path] [--transpose] [--sparse]

With no --key, list the candidate 2-D datasets (reuse adata ls logic) and exit.

Delimited text (read_csv, read_text, .tsv/.txt/.tab, optionally gzipped)

Dense matrix, first row = column names, first column = row names.

  • Parse row blocks as a stream; never load the file.
  • --sparse writes CSR via sparsify from core/convert.py (recommended for count tables).
  • --delimiter, --first-column-names/--no-first-column-names, --transpose (many tables are genes × cells).
  • Needs a tiny CSV reader without pandas — share it with Convert: Parse Biosciences split-pipe outputs #30.

Out of scope

Excel, .soft.gz, UMI-tools — listed in #31.

Tests

Compare with scanpy.read_hdf / read_csv / read_text; perf guard on row-block parsing.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions