Skip to content
View eduardogade's full-sized avatar

Block or report eduardogade

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
eduardogade/README.md

Typing SVG

    Knowledge Domains | Data Engineering & Platform Systems across cloud, hybrid, and regulated data environments

    Technical Quality | Focused on production-grade pipelines: from raw ingestion -> trusted datasets -> analytics-ready products

    Engineering Practice | Strong emphasis on reliability, reproducibility, auditability, and systems that age well

    Broad Curriculum | Experience spanning ETL/ELT, distributed processing, cloud/HPC, data quality, and ML-adjacent systems

I would like to know more...

Hello, and welcome to my profile. My name is Eduardo - grab a cup of coffee and allow me to introduce myself.

I build data platforms and production-grade data systems where messy real-world inputs are transformed into reliable, auditable, and analysis-ready products.

In practical terms, my work lives somewhere between:

  • Data ingestion - where reality arrives poorly formatted and with opinions.
  • Transformation layers - where Python, SQL, Spark, and modeling discipline try to restore civilization.
  • Data quality - because a pipeline that runs and a pipeline that is correct are not the same animal.
  • Delivery - where analytics, BI, ML, and internal users need datasets that are stable enough to trust and boring enough to maintain.

I specialize in Python, SQL, PySpark/Spark, cloud and hybrid data infrastructure, ETL/ELT pipelines, data modeling, validation frameworks, and reproducible platform workflows. I have worked across environments involving AWS, GCP, Azure, HPC clusters, Docker/Kubernetes, CI/CD, and large heterogeneous datasets that rarely introduce themselves politely.

My engineering bias is simple: build systems that are clear, testable, observable, and maintainable after the original excitement has left the room.

This usually translates to:

  • Production data pipelines with strong reproducibility, monitoring, and failure handling
  • Scalable batch and distributed processing for high-volume analytical workloads
  • Data modeling layers supporting analytics, BI, reporting, and ML workflows
  • Validation and governance practices for environments where correctness is not decorative
  • Cloud, hybrid, and HPC workflows designed to survive both scale and human memory

My background includes complex healthcare and enterprise data environments, but my current professional identity is straightforward: Senior Data Engineer / Data Platform Engineer. Machine Learning is still part of the toolbox, but the main job is now the plumbing, contracts, orchestration, and reliability that make downstream intelligence possible.

I value clean design, explicit trade-offs, and systems that are understandable by humans - not just machines with suspicious confidence.

Ethics, reproducibility, and long-term sustainability are not optional; they are part of the job.

Availability | Currently open to remote, hybrid, relocation-friendly, and long-term Data Engineering / Data Platform roles. See contact details. Relocation and onboarding take planning - good systems (and good moves) benefit from doing things properly.

Cheers.


    2025 | Committer | Awarded Apache Spark Committer Status | The Apache Software Foundation (ASF) | Finland & Brazil

    2022 | Senior Transition | Senior Data Engineer | Turku Biosciences & Brazilian Ministry of Health | Finland & Brazil

    2020 | Outreach | Award-Winning COVID-19 Outreach Campaign | Göttingen General Hospital | Germany

    2017 | Patent | LAG3-Targeting Cancer Therapy | Current owner: Bristol Myers Squibb | USA

    2016 | Industry Transition | Data Engineer | Dana-Farber Cancer Institute | USA

    2013 | Research | Computational Biology Researcher | RWTH University Medical School | Germany

I would like to know more...

Career Analytics

KEY MILESTONES

  • Data Engineer

  • Data & Health Researcher

  • Diplomat between stakeholders

    • Advanced storytelling techniques
    • Translate real-world business problems into systems
    • Sharp and operational communication
  • Current: Cloud & Platform Systems

    • Developing efficient cloud-based ecosystems
    • Managing 4 DEs, 2 Data Scientists and 1 Dev
    • Filed 2 patents and improved operating margin by ~18%

WRITER AND EDUCATOR


name: "Eduardo Gusmao"
role: "Senior Data Engineer | Data Platform Engineer"
contact: "Recife, Brazil | eduardogade@gmail.com | github.com/eduardogade | linkedin.com/in/eduardogade"
languages: "English fluent | Portuguese native | Spanish B2 | German A2 | Finnish A1"

education: "2x PhD in Biomedical Informatics and Data Engineering / Computational Life Sciences; BSc + MSc in Computer Science"

summary: "Data Platform Engineer with 8+ years designing scalable data platforms, distributed systems, and production-grade pipelines across healthcare, life sciences, and enterprise environments. Strong Python, SQL, PySpark/Spark, AWS, Docker/Kubernetes, CI/CD, data quality, and analytics/BI platform experience."

professional_engagements:
  current_role:
    company: "Turku Biosciences / Brazilian Ministry of Health"
    title: "Senior Data Engineer"
    location: "Finland / Brazil"
    date: "Sep 2022 - Present"
    scope:
      - "Lead a national-scale precision-medicine data platform integrating genomic, phenotypic, and clinical EHR data for 65,000+ individuals."
      - "Build scalable Python, SQL, PySpark/Spark, Databricks-adjacent, HPC/SLURM, and API-driven ingestion and transformation workflows."
      - "Deliver regulated ingestion, validation, governance, PII-compliant processing, observability, idempotency, and data quality controls."
      - "Support analytics, BI, and ML workloads through reusable integration layers, backend data services, and optimized Parquet-based processing."

development_environment:
  infrastructure: "AWS | Azure | GCP | HPC/SLURM | Docker | Kubernetes | Terraform | GitHub Actions"
  languages: "Python | SQL | PySpark | Bash/Shell | Scala | Java | C/C++ | YAML | HCL"
  data_stack: "PySpark | Spark | Pandas | Polars | NumPy | BigQuery | PostgreSQL | Parquet | JSON | dbt | dimensional modeling"
  platform_engineering: "ETL/ELT | ingestion frameworks | transformation layers | platform APIs | CI/CD | observability | validation | data quality"
  ml_ai_stack: "ML pipelines | MLOps | feature engineering | LLM APIs | embeddings | RAG | Hugging Face"
  collaboration: "Agile/Scrum | stakeholder enablement | analytics teams | data scientists | engineers | product | infrastructure | security"
I would like to know more...
name: "Eduardo Gusmao"
role: "Senior Data Engineer | Data Platform Engineer | Cloud Data Engineer"
location: "Recife, Brazil"
contact: "eduardogade@gmail.com | github.com/eduardogade | linkedin.com/in/eduardogade"
languages: "English fluent | Portuguese native | Spanish B2 | German A2 | Finnish A1"

summary: "Data Platform Engineer with 8+ years of experience designing and operating scalable data platforms, distributed data systems, and production-grade pipelines across healthcare, life sciences, and enterprise environments. Strong expertise in Python and SQL, with hands-on experience in PySpark/Spark, Kafka-style event-driven workflows, Airflow, AWS, Docker/Kubernetes, Terraform, and CI/CD to build reliable data infrastructure and platform services."

core_expertise:
  - "Data Engineering"
  - "Data Platform Engineering"
  - "Cloud and Hybrid Data Infrastructure"
  - "Distributed Data Systems"
  - "ETL/ELT Pipelines"
  - "Data Modeling and Analytics Engineering"
  - "Data Quality and Governance"
  - "Healthcare and Life Sciences Data"
  - "Machine Learning Data Pipelines"
  - "Production Reliability and Observability"

career_profile:
  - "8+ years building scalable data platforms and distributed data systems"
  - "Production-grade ingestion, transformation, validation, observability, and internal tooling"
  - "Strong Python, SQL, PySpark/Spark, cloud, CI/CD, Docker/Kubernetes, and data quality background"
  - "Experience supporting analytics, BI, ML workflows, and mission-critical data products"
  - "Comfortable translating complex stakeholder requirements into maintainable platform capabilities"

professional_engagements:
  current:
    company: "Turku Biosciences / Brazilian Ministry of Health"
    title: "Senior Data Engineer"
    location: "Finland / Brazil"
    date: "Sep 2022 - Present"
    scope:
      - "Lead the design and delivery of a national-scale data platform for precision medicine, integrating multi-modal genomic, phenotypic, and clinical EHR data for 65,000+ individuals."
      - "Build scalable Python-based pipelines, distributed systems, PySpark/Spark workflows, Databricks-adjacent processing, HPC/SLURM execution, and API-driven ingestion workflows."
      - "Architect production-grade data platform services for regulated ingestion, validation, transformation, governance, PII-compliant processing, reproducible workflows, and quality/reliability controls."
      - "Develop high-performance processing and modeling layers using Python, SQL, PySpark, partitioning strategies, Parquet formats, and distributed query tuning."
      - "Design reusable integration layers and backend data services connecting heterogeneous clinical, genomic, and ERP/SAP data sources."
      - "Enable event-driven workflows, orchestration patterns, and batch/streaming-adjacent pipelines supporting analytics, BI, and ML systems."
      - "Collaborate with product, analytics, engineering, and infrastructure stakeholders to deliver platform capabilities, CI/CD, Docker/Kubernetes workloads, observability, logging, alerting, and performance tuning."
    outcomes:
      - "Integrated precision-medicine datasets for 65,000+ individuals."
      - "Improved pipeline efficiency by approximately 25%."
      - "Reduced storage costs by approximately 80%."
      - "Enabled more than 40% faster data delivery for downstream analytics, BI, and ML systems."

  previous_mid:
    company: "Göttingen General Hospital"
    title: "Data Engineer II"
    location: "Germany"
    date: "Mar 2019 - Sep 2022"
    scope:
      - "Designed and implemented scalable data platform services using Python and SQL on cloud and hybrid environments."
      - "Enabled reliable ingestion, transformation, and low-latency access for downstream analytics, BI, and application workloads."
      - "Developed and optimized high-performance ETL/ELT pipelines with Python and PySpark, leveraging distributed processing, batch workflows, and orchestration patterns."
      - "Refactored legacy systems into modular, production-grade platform services with CI/CD, automated testing, monitoring/logging, idempotency, retries, and robust error handling."
      - "Built reusable data processing frameworks and integration layers for large-scale heterogeneous datasets."
      - "Applied data modeling, validation, lifecycle standards, and governance across 12 cross-functional teams in a distributed environment."
    outcomes:
      - "Improved data availability and system responsiveness by approximately 33%."
      - "Improved reliability, maintainability, and operational efficiency by approximately 50-60%."
      - "Supported consistent data lifecycle practices across 12 cross-functional teams."

  previous_old:
    company: "Dana-Farber Cancer Institute"
    title: "Data Engineer I"
    location: "USA"
    date: "Jan 2016 - Mar 2019"
    scope:
      - "Developed cloud-native data platform services supporting large-scale drug discovery."
      - "Built Python-based ETL/ELT pipelines and API-driven integration layers for heterogeneous biomedical, operational, and financial datasets."
      - "Implemented end-to-end data processing pipelines using Python, SQL, and PySpark on Apache Spark distributed systems."
      - "Enabled scalable ingestion, transformation, validation, and batch workflows for analytics and ML-driven applications."
      - "Collaborated with product, analytics, and research stakeholders to define KPIs and translate requirements into data models, backend data logic, and reusable platform components."
      - "Contributed to production-grade data engineering practices including Git version control, validation checks, documentation, maintainable system design, reliability, and reproducibility."
    outcomes:
      - "Improved data accessibility and reduced operational costs by more than 25%."
      - "Supported analytics and ML-driven applications through reusable data platform components."
      - "Established reliable, reproducible lifecycle standards for heterogeneous biomedical and operational data."

education:
  phd_biomedical_informatics:
    degree: "Ph.D. in Biomedical Informatics"
    institution: "Harvard Medical School"
    location: "Boston / Cambridge, USA"
    date: "2013 - 2017"

  phd_computational_life_sciences:
    degree: "Ph.D. Dr. rer. nat. in Data Engineering and Computational Life Sciences"
    institution: "RWTH Aachen University"
    location: "Aachen, Germany"
    date: "2011 - 2015"

  bachelor_master_computer_science:
    degree: "B.Sc. and M.Sc. in Computer Science"
    institution: "Federal University of Pernambuco"
    location: "Recife, Brazil"
    date: "2008 - 2011"

technical_strengths:
  programming:
    primary: ["Python", "SQL", "PySpark", "Spark SQL", "Bash/Shell"]
    secondary: ["Scala", "Java", "C/C++", "YAML", "HCL"]
    concepts: ["REST APIs", "Async programming", "Data serialization", "Production-grade software engineering", "Parquet", "JSON"]

  data_platform_engineering:
    capabilities: ["Scalable data platforms", "Distributed data systems", "Internal data tooling", "Reusable ingestion frameworks", "Transformation layers", "Platform APIs", "Developer-facing abstractions", "Self-service data capabilities", "Analytics enablement", "ML workflow support", "BI workload support"]

  distributed_data_systems:
    tools: ["PySpark", "Apache Spark", "Pandas", "Polars", "NumPy"]
    capabilities: ["Large-scale processing", "Distributed compute", "Performance tuning", "Partitioning", "Query optimization", "Resource efficiency", "Batch pipelines", "Streaming-adjacent pipelines", "Kafka", "Spark Streaming patterns"]

  cloud_hybrid_infrastructure:
    cloud: ["AWS", "Azure", "GCP"]
    infrastructure: ["HPC/SLURM", "Docker", "Kubernetes", "Terraform", "GitHub Actions", "GitLab CI"]
    capabilities: ["Cloud-native data infrastructure", "Hybrid data infrastructure", "Infrastructure-aware engineering", "Containerized workloads", "Deployment environments", "Scalable platform operations"]

  hadoop_on_prem_ecosystems:
    technologies: ["HDFS", "YARN", "Hive", "Kerberos"]
    capabilities: ["Distributed storage patterns", "Distributed compute patterns", "Legacy-to-modern platform evolution", "Secure access-controlled data environments"]

  data_modeling_tooling:
   : ["Dimensional modeling", "Semantic modeling", "Schema design", "Metadata management", "Transformation layers", "Data contracts", "Lineage", "Modeling standards", "Analytics enablement", "Platform consistency"]
    tools: ["dbt", "BigQuery", "PostgreSQL", "Parquet", "JSON"]

  software_engineering_devops_reliability:
    tools: ["Git", "GitHub", "GitHub Actions", "GitLab CI", "Docker", "Kubernetes", "Terraform"]
    practices: ["CI/CD pipelines", "Automated testing", "Deployment automation", "Monitoring", "Logging", "Alerting", "Observability", "Incident response", "Idempotency", "Retries", "SLA/SLO thinking", "Fault-tolerant design"]

  machine_learning_data_pipelines:
    capabilities: ["ML pipelines", "MLOps", "Feature engineering", "Data preparation", "Personalization workflows", "AI-enabled data workflows", "Production-oriented ML data support"]
    ai_llm: ["LLM APIs", "Embedding pipelines", "RAG", "Hugging Face"]

  data_security_governance_quality:
    capabilities: ["Data privacy", "PII-aware processing", "Compliance-aware pipelines", "Access control", "Validation strategies", "Auditability", "Data quality checks", "Governance practices", "Secure data lifecycle management", "Reliability controls", "Consistency checks"]

  processes_collaboration:
    practices: ["High-ownership engineering mindset", "Agile/Scrum", "Cross-functional collaboration", "Stakeholder enablement", "Requirements translation", "Technical documentation", "Platform capability delivery"]
    collaborators: ["Analysts", "Data scientists", "Engineers", "Product teams", "Infrastructure teams", "Security teams"]

development_environment:
  hardware: ["Apple Silicon", "ARM", "Intel", "NVIDIA GPU environments", "HPC clusters"]

  operating_systems: ["macOS", "Ubuntu", "Debian", "Fedora", "Windows"]

  infrastructure:
    cloud_computing: ["AWS", "Azure", "GCP"]
    hpc: ["SLURM", "OpenPBS", "Distributed compute environments"]
    containers: ["Docker", "Kubernetes", "Singularity"]
    infrastructure_as_code: ["Terraform", "HCL", "Cloud deployment"]

  languages:
    data_engineering: ["Python", "SQL", "PySpark", "Spark SQL", "Bash/Shell"]
    systems_and_general: ["C/C++", "Java", "Scala"]
    markup_and_config: ["YAML", "Markdown", "LaTeX", "HTML/CSS", "HCL"]

  data_stack:
    distributed_processing: ["Apache Spark", "PySpark", "Spark SQL", "Pandas", "Polars", "NumPy"]
    storage_formats: ["Parquet", "JSON", "CSV", "HDF5"]
    databases_and_warehouses: ["BigQuery", "PostgreSQL", "MongoDB", "DynamoDB", "Relational databases", "NoSQL databases"]
    modeling_and_quality: ["Dimensional modeling", "Semantic modeling", "Schema design", "Data contracts", "Lineage", "Validation checks", "Data quality checks", "dbt"]

  ml_ai_stack:
    frameworks_and_tools: ["PyTorch", "TensorFlow", "Keras", "Scikit-Learn", "Hugging Face", "NLTK"]
    workflows: ["ML pipelines", "MLOps", "Feature engineering", "Embedding pipelines", "RAG", "LLM APIs"]

  systems_tooling:
    version_control: ["Git", "GitHub"]
    packaging_and_environments: ["pip", "poetry", "micromamba", "mamba", "conda", "npm"]
    ci_cd: ["GitHub Actions", "GitLab CI"]
    observability: ["Logging", "Monitoring", "Alerting", "Observability", "Prometheus", "Grafana"]

github_positioning:
  short_pitch: "I build reliable data platforms, distributed pipelines, and production-ready data systems for analytics, BI, ML, and healthcare/life-sciences workloads."
  engineering_style:
    - "Clean, maintainable, typed Python"
    - "Data quality and reliability first"
    - "Production-aware platform design"
    - "Reproducible workflows"
    - "Strong documentation"
    - "Pragmatic cloud and hybrid infrastructure"

    Email | eduardo@gusmaolab.org

    LinkedIn | https://www.linkedin.com/in/eduardogade/

    Location | Recife, Brazil | Remote-friendly

    Status | Open to Data Engineering roles

I would like to know more...

Professional Profiles

    LinkedIn: https://www.linkedin.com/in/eduardogade/

    Website & Blog: https://www.gusmaolab.org

    CV/Resume: https://www.gusmaolab.org/cv/CV_Eduardo_Gusmao.pdf

    Stack Overflow: https://stackoverflow.com/users/32223943/eduardo-gusmao

    Medium: https://medium.com/@eduardogade

Practical notes

    Preferred contact: Email | LinkedIn

    Response time: 1-2 business days

    Open to remote, hybrid, or relocation

Details

    See [availability & engagement details](#availability)

    See [writing & communication details](#communication)

    See [education](#education) & [leadership details](#career)


Apollo

Apollo

A unified suite of post-hoc statistical procedures with bias-aware corrections designed for metrics common in computational and ML/DL pipelines.

View Repository →
Blacksmith

Blacksmith

A high-performance genotype analysis framework for streamlined quality control, variant graph construction, and interactive network visualization.

View Repository →
Olympus

Olympus

A unified framework for discovering, analyzing, integrating, and visualizing regulatory motifs and transcription factor binding sites across bulk, single-cell, and long-read sequencing modalities.

View Repository →
I would like to know more...
Apollo

Apollo

A unified suite of post-hoc statistical procedures with bias-aware corrections designed for metrics common in computational and ML/DL pipelines.
Blacksmith

Blacksmith

A high-performance genotype analysis framework for streamlined quality control, variant graph construction, and interactive network visualization.
Olympus

Olympus

A unified framework for discovering, analyzing, integrating, and visualizing regulatory motifs and transcription factor binding sites across bulk, single-cell, and long-read sequencing modalities.
Bloom

Bloom

A framework for chromatin architecture data processing, handling and analysis.
Musique

Musique

A unified transcriptomics analysis framework supporting bulk, single-cell, long-read, short-read, and spatial expression workflows with integrated normalization, quantification, modeling, and visualization tools.
Wildlife

Wildlife

A unified deep learning framework for high-performance multimodal data imputation, integrating neural operators for tabular, EHR, imaging, audio, video, and biological datasets.
Fabric

Fabric

A collection of health informatics algorithms and tools.
Uqbar

Uqbar

Ubiquitously Broad Automation and Architecture — a collection of tools for small task automation.
GusmaoLab

GusmaoLab

Source of Eduardo Gusmao's lab website — portfolio, technical blog, and public-facing documentation.

    Apache Software Foundation | Committer | Apache Spark | 2025 - Present

    Maintainer | 9+ actively maintained research and engineering tools | See Pinned Repositories

    Community & Volunteering | Global Burden of Disease, TransEmpregos, ABRATA | See Causes

    Status | Curating additional public contribution history

I would like to know more...

Open-source involvement is treated as an extension of day-to-day engineering practice rather than a separate résumé line - it is where tooling, methods, and lessons from production work get generalized and given back.

The most concrete example is the Apache Spark committer status awarded by the Apache Software Foundation in 2025, which followed sustained engagement with the project's distributed-processing internals through professional and personal work. Beyond that, most of the ongoing contribution activity currently lives in the personal toolkit showcased under Pinned Repositories - actively maintained libraries for statistics, regulatory genomics, chromatin analysis, and small-scale automation - and in the volunteering commitments described under Causes.

A broader, evidence-backed view of pull requests, issue triage, and cross-project reviews across external organizations is still being organized for public presentation. Rather than publish an incomplete or padded picture, this subsection will be expanded once that material is ready to stand on its own.



3D contribution graph

I would like to know more...

Additional breakdowns of repositories, commit language distribution, and contribution timing - not already shown above.


contribution snake


This section is reserved for repository-influence and collaboration analytics - dependency relationships, external review activity, and community engagement across the tools in Pinned Repositories.

I would like to know more...

The dashboard above already covers language mix, commit activity, and contribution cadence well. What is intentionally not here yet is a second layer of analytics: how these tools get used and reviewed outside of direct authorship - dependency graphs, downstream adoption, and external review activity.

No currently available public tooling produces that view at a quality bar consistent with the rest of this page without either being self-hosted or relying on metrics (like early-stage star counts) that would be more decorative than informative for research tooling with a narrow, specialist audience. Rather than fill the space with a vanity widget, this section stays intentionally reserved until metrics that are actually meaningful - closer to genuine downstream engineering usage than surface-level GitHub counters - become available.


apple
Apple
python
Python
pytorch
PyTorch
gatk
GATK
git
Git
snakemake
Snakemake
gradio
Gradio
docker
Docker
aws
AWS
jira
Jira
linux
Linux
r
R
tensorflow
TensorFlow
bioconductor
Bioconductor
github
GitHub
nextflow
Nextflow
fastapi
FastAPI
kubernetes
K8s
terraform
Terraform
grafana
Grafana
vscode
VsCode
bash
Bash
jax
JAX
ruff
Ruff
githubactions
GActions
Mamba/Conda
Mamba
postgresql
Postgres
redis
Redis
databricks.svg
DtBricks
prometheus
Prometheus
I would like to know more...

Complete technology inventory, synchronized with CV.yaml and grouped by domain. Icons shown where a stable public icon set provides one; unmarked names are listed as text only.

Programming Languages

Python SQL PySpark/Spark SQL Scala Java C/C++ Bash/Shell Go Rust Julia Kotlin TypeScript JavaScript C# Ruby PHP HCL

Data Engineering & Orchestration

ETL/ELT dbt Airflow Kafka Dimensional Modeling Schema Design Data Contracts Lineage Data Quality

Databases & Storage

PostgreSQL MySQL MongoDB Redis BigQuery DynamoDB ArangoDB Pinecone FAISS Parquet/JSON/CSV HDF5/Zarr

Cloud & Infrastructure

AWS Azure GCP Terraform CloudFormation Pulumi Docker Kubernetes Helm Singularity Podman SLURM/OpenPBS/MPI GitHub Actions GitLab CI Jenkins

Machine Learning & AI

PyTorch TensorFlow JAX Keras Scikit-Learn Hugging Face NLTK Ray TensorRT LLM APIs / RAG / Embeddings MLOps

Scientific & Bioinformatics Computing

Pandas NumPy SciPy Polars Dask Bioconductor PySAM GATK Snakemake Nextflow OpenCV PyCaret

Web, APIs & Dashboards

FastAPI Django React Next.js Express.js Gradio Streamlit Dash REST GraphQL gRPC Power BI

Observability & DevOps

Prometheus Grafana Ruff Git pip/poetry/(micro)mamba

Documentation & Markup

YAML Markdown LaTeX HTML/CSS Quarto


    Machine & Deep Learning | Repository | Publication

    Variational Inference | Repository | Publication

    Precision Medicine | Repository | Publication

    Regulatory Genomics | Repository | Publication

I would like to know more...

Selected Publications (decreasing order by year)

Global age-sex-specific all-cause mortality and life expectancy estimates for 204 countries and territories and 660 subnational locations, 1950-2023: a demographic analysis for the Global Burden of Disease Study 2023

The Lancet · Oct 18, 2025

Contributions:

  • Responsible for orchestrating the LATAM-branch with 45+ PIs and 200+ researchers.
  • Horizontal meetings for data and experience sharing have shown great success, with ~380% more efficiency than the second most efficient branch - per capita.
  • Has solved pharmacological conflict of interests by cross-deployment and blind-genotype blind-phenotype strategy, which exhibit 17% increased accuracy over North America (first COI - percapita) and 5% over Asia (second COI - per capita).

A ONECUT1 regulatory, non-coding region in pancreatic development and diabetes

Cell Reports · Nov 26, 2024

Contributions:

  • The tool Bloom has increased analysis mechanism by promoting different views into the regulatory spatial configuration, resulting in ~50% wet-lab equipment cost reduction and solving a stalled-case.
  • Provided personal guidance towards architecture and Hi-C methodology, saving 15% overall lab-time.
  • Overall, this was the first non-trivial non-intermediary-distance (>1Gbp) lncRNA interference in a region unknown to be a regulatory enhancer.

Global, regional, and national burden of diabetes from 1990 to 2021, with projections of prevalence to 2050: a systematic analysis for the Global Burden of Disease Study 2021

The Lancet · Jul 15, 2023

Contributions:

  • Responsible for orchestrating a team of 3 brazilian PIs and 5 independent investigators.
  • Used scrum, coupled with CRISP-DM, delivering net gains (profitability converted back) through network revenue saving and wet/dry-lab material cost reduction.
  • Developed national-scale geno/phenotype QC pipeline - Fabric (Phenoteka Module) - used across 20+ institutes.

100,000 Genomes Pilot on Rare-Disease Diagnosis in Health Care - Preliminary Report

The New England Journal of Medicine · Oct 11, 2022

Contributions:

  • Developed Blacksmith, that coupled with Bloom improved operating margin by over 15%.
  • Freed at least 15 engineering hours per week with Blacksmith coupled with Apollo.
  • Intending to lower carbon footprint, we adopted a trademarked DB 'bit-brushing' methodology (currently owned by Databricks Inc.).

HMGB1 coordinates SASP‐related chromatin folding and RNA homeostasis on the path to senescence

Molecular Systems Biology · Jun 24, 2021

Contributions:

  • Analized Spatial Chromatin Biology and RNA-seq to identify - for the first time - HMGB1 as a 'rheostat' factor.
  • Reduced cloud compute costs by 40% using Apollo's strong mathematical features and Bloom to analyse Chromatin conformation.
  • After this project's results, we have earned an ESG compliance through impeccable waste management and safety handling.

Redundant and specific roles of cohesin STAG subunits in chromatin looping and transcriptional control

Genome Research · Apr 6, 2020

Contributions:

  • Analized most omics in a single project: ChIP-seq, degron-X, RNA-seq, Hi-C, STORM, DNase-seq, ATAC-seq, MSMS and MS-based microscopy.
  • Developed Musique, shortening development cycles by ~9 weeks.
  • Musique saved 300 GPU-hours per month by performing simple heuristics which are generalizeable to any dataset.

Spatial chromosome folding and active transcription drive DNA fragility and formation of oncogenic MLL translocations

Molecular Cell · Jul 25, 2019

Contributions:

  • Patented technique for BLISS-seq data processing, earning ~25% extra funds for the laboratory.
  • Lower wet-lab costs using dry-lab tools by ~30% (estimated for this project); achieving reproducible and insightful results on MLL fusions.
  • Created the triple-correlation method. Translating category theory into a real-world phenomenon.

HMGB2 loss upon senescence entry disrupts genomic organisation and induces CTCF clustering across cell types

Molecular Cell · May 17, 2018

Contributions:

  • Developed Bloom and Apollo, which reduced processing time by at least 3 months.
  • Very agile methodology with microprocessed multicycled days, leading to novel discoveries and decreasing overall time-to-delivery.
  • Reduced local infrastructure storage footprint by ~100TB with Bloom & Apollo.

Integrated genomic and molecular characterization of cervical cancer

Nature · Jan 23, 2017

Contributions:

  • Devised bioinformatics pipelines with collaborators and created the Gaussian-as-DPMM method of clustering, increasing speed by, at least, ~100x.
  • Clustering was able to identify 3+ unique subtypes never previously reported.
  • Created a deep regulatory network, especially with SHKBP1 ERBB3 and TGFBR2; which contained 98% of the cancer mortality information variability.

Analysis of computational footprinting methods for DNase sequencing experiments

Nature Methods · Feb 22, 2016

Contributions:

  • Landmark study on comparing 12+ footprinting methods. The study was the cover of Nature Methods magazine.
  • Without any dry experiment, we were able to identify the limits of sequencing technologies, and propose results that exceded ~5% AUPR of known methods.
  • Our method - Olympus (published in 2023) - offers ~7x most complete analysis of regulatory genomics than any other tool.

Epigenetic program and transcription factor circuitry of dendritic cell development

Nucleic Acids Research · Oct 17, 2015

Contributions:

  • First use of Faun, the motif enrichment analysis that uses hypergeometric distributions to query the sensitivity and specificity of TF occupancy in a certain genomic region.
  • Proposed the usage of Cytoscape, widely minimising meeting preparation time by ~25%.
  • Proposed use of fewer histone modification essays by recreating chromatin states in silico; thus, minimizing project costs by ~30%

Blog Atlas Learning Series Medium Book

I would like to know more...

I treat writing as an engineering discipline, not an afterthought bolted onto the end of a project. A pipeline that nobody can explain six months later is not really finished, no matter how well it runs today - the explanation is part of the deliverable, and it degrades just like code does if nobody maintains it.

Most of what I write falls into two categories: documentation that has to survive contact with someone else's Monday morning, and architecture notes that have to survive contact with my own memory a year from now. Both audiences are unforgiving in the same way - they punish vagueness, reward precision, and do not care how elegant the underlying system is if the explanation doesn't hold up. Writing clearly about a system is often the fastest way to discover that it isn't as clearly designed as I thought.

The Atlas Learning Series and the posts on Gusmao Lab and Medium exist for the same reason internal design docs exist: complexity that stays in one person's head is a liability, and complexity that gets written down - honestly, without inflating it - becomes something a team can actually build on.


    Status | Open to new opportunities | Data Engineering / Data Platform roles

    Work style | Remote, hybrid, or on-site | Relocation-friendly

    Geography | Europe → Brazil → Worldwide, roughly in that order of preference

    Engagement | Long-term, technically demanding roles preferred over short-term contracts

I would like to know more...

Geography is a genuine preference, not a hard boundary. Europe and Brazil are where my professional and personal ties are strongest, so a role based there - or a remote/hybrid arrangement rooted there - tends to be the smoothest fit. That said, a worldwide, well-run, technically serious opportunity is always worth a conversation; relocation and onboarding just take a bit more planning, the same way any well-run migration does.

What I optimize for in a team is closer to craftsmanship than throughput. I would rather ship a system that is a little slower to build and considerably easier to operate, extend, and hand off, than one that is fast to demo and expensive to live with. That preference shows up as a bias toward explicit trade-offs, readable code, tests that actually catch regressions, and documentation that survives the person who wrote it moving on to something else.

Productivity, in my experience, tracks environment more than effort. Teams with reproducible builds, sane CI/CD, infrastructure-as-code, and automation around the boring parts of the job consistently ship better software than teams without them - independent of talent, tooling brand, or operating system. I look for that kind of environment because it is where careful engineering actually compounds instead of getting rebuilt from scratch every eighteen months, and I try to leave every team I join a little closer to it than I found it.

Collaboration matters just as much as any individual technical choice. The best systems I have worked on came out of teams that argued about trade-offs early, wrote decisions down, and treated architecture as something owned collectively rather than defended individually - that is the kind of team I look for, and the kind I try to build.


Roadmap illustration placeholder

This section will eventually map out long-term technical direction rather than a short-term task list: platform evolution, research directions worth pursuing further, open-source initiatives planned for the tools under Pinned Repositories, and the engineering initiatives that connect them. The illustration above is a placeholder reserved for that map.

I would like to know more...

Roadmap illustration placeholder, expanded

The expanded version of this section will lay out where the underlying platforms and research directions are headed over the next several years - which parts of the current toolkit graduate from personal research tools into broader shared infrastructure, which open questions from the publication record are worth a dedicated engineering push, and which engineering practices (observability, reproducibility, data contracts) are being generalized across projects rather than reinvented per repository. The illustration above will be replaced with the final roadmap artwork once that direction is settled.


    2017 | PhD | Biomedical Informatics | Harvard Medical School

    2015 | PhD | Life Sciences | RWTH Aachen University

    2011 | MSc | Machine & Deep Learning | Federal University of Pernambuco

    2008 | BSc | Computer Science | Federal University of Pernambuco

I would like to know more...
Ph.D. in Biomedical Informatics — Harvard Medical School (2013 - 2017)

Doctoral work carried out in parallel with a data engineering role at the Dana-Farber Cancer Institute, focused on computational pipelines supporting CRISPR-based immunotherapy target discovery. Contributed to target-discovery work that advanced into clinical development for LAG-3-based immunotherapy, later commercialized as Opdualag (Relatlimab + Nivolumab). Co-authored studies published during this period spanning cervical cancer genomics (Nature), chromatin organization in senescence (Molecular Cell), and computational footprinting methods for regulatory genomics (Nature Methods) - see Research.

Ph.D. Dr. rer. nat. in Data Engineering and Computational Life Sciences — RWTH Aachen University (2011 - 2015)

Doctoral research in regulatory genomics and epigenomics, centered on computational footprinting methods for DNase-seq and related chromatin assays. This period produced the analytical groundwork later published as a landmark comparison of 12+ footprinting methods (cover article, Nature Methods) and the method that became Olympus, alongside early chromatin-architecture work that would grow into Bloom.

M.Sc. in Computer Science (Machine & Deep Learning) — Federal University of Pernambuco (2010 - 2011)

This subsection will summarize the thesis focus, key coursework, and any early research output from the Master's program once that material is organized for public presentation.

B.Sc. in Computer Science — Federal University of Pernambuco (2008 - 2011)

This subsection will summarize notable coursework, early projects, and foundational technical milestones from the undergraduate program once that material is organized for public presentation.


    Craftsmanship | Software is read far more often than it is written; optimize accordingly

    Reproducibility | If a result cannot be rebuilt from scratch, it is not yet a result

    Documentation & Testing | Explicit beats implicit; tested beats assumed

    Data Governance | Especially with healthcare and other sensitive data, privacy and auditability are not optional

    Scientific Integrity | Reproducible engineering practices are how research claims earn trust

I would like to know more...

Most of my engineering principles come from having worked on data that describes real people - patients, cohorts, research participants - long before it became rows in a table. That background makes certain things non-negotiable: sensitive data gets handled with explicit access controls and audit trails, not "we'll get to it later" promises, and a pipeline that cannot explain its own provenance is not trustworthy no matter how good its output looks.

Maintainability is the other constant. Code that only its author can safely change is a liability with a delay timer on it, and documentation that describes what a system does two versions ago is worse than no documentation at all, because it actively misleads. I would rather spend an extra afternoon on a clear interface and an honest README than leave that debt for whoever inherits the system - often enough, that person is me, eighteen months later, with no memory of the clever shortcut I took.

Scientific reproducibility and software engineering discipline turned out to be the same skill wearing different clothes. A result that cannot be rebuilt from raw inputs by someone else is not really a result yet, just a claim; the same containerized, version-controlled, tested rigor that makes a data pipeline trustworthy in production is what makes a research finding trustworthy in a paper. I try not to draw a hard line between the two.

A few thinkers show up in how I approach this, more as working tools than as heroes. Wittgenstein's attention to the limits of language maps directly onto naming things well and writing documentation that means exactly what it says - ambiguity in a spec is a bug, not a style choice. Camus's insistence on working honestly inside an absurd, uncertain world is a fair description of what building on messy real-world data actually feels like: you do not get certainty, you get discipline. And the habit of asking who actually benefits from a given system, and who quietly bears its costs, is a useful lens for thinking about the social impact of the technology I help build - not a political position, just a question worth asking before shipping.

Flagship: 🏳️‍⚧️ | 🏳️‍🌈 | 🇺🇳


Global Burden of Disease

Collaborator · 2018 - Present

Contributing to one of the world's largest collaborative efforts for measuring disease burden through data infrastructure, engineering, and scientific collaboration.
TransEmpregos

Volunteer · 2024 - Present

Supporting infrastructure and data engineering for employability initiatives focused on transgender professionals.
ABRATA

Volunteer · 2022 - 2025

Supporting initiatives that improve awareness and education regarding mood disorders and mental health.
I would like to know more...

Global Burden of Disease is one of the largest ongoing scientific collaborations in public health, coordinating thousands of researchers across nearly every country to produce comparable estimates of disease, injury, and risk-factor burden over time. For a data engineer, it is also a genuinely hard distributed-collaboration problem - reconciling heterogeneous data sources, methods, and reporting standards across institutions that rarely share infrastructure. Participating in that network connects day-to-day pipeline work to public-health decisions made well beyond any single hospital or country.

TransEmpregos works to close the employment gap faced by transgender professionals in Brazil, one of the countries where that gap is widest. The technical need is unglamorous and familiar - reliable data infrastructure so the organization can run, match opportunities, and measure impact - but the outcome is concrete: real people getting a fairer shot at stable employment.

ABRATA (Associação Brasileira de Familiares, Amigos e Portadores de Transtornos Afetivos) works on public awareness and education around mood disorders in Brazil, where mental-health literacy still lags far behind physical-health literacy. Supporting an organization like this is a reminder that not every meaningful contribution has to be technical - sometimes it is simply showing up.


Powerlifting placeholder

Powerlifting

This card is reserved for the story of how powerlifting became part of a long-term routine - the discipline of programmed progress, patience with plateaus, and the same respect for fundamentals that shows up in engineering work. Placeholder narrative pending the owner's personal write-up.

I would like to know more...
Open source contributor placeholder

Open Source Contributor

This card is reserved for a closer look at the personal research toolkit maintained under Pinned Repositories and the broader open-source habits behind it. Placeholder narrative pending the owner's personal write-up.

Lecturer and public speaker placeholder

Lecturer & Public Speaker

This card is reserved for a closer look at teaching CS at UFPE/CIn and TUM/SLS, keynote talks at NeurIPS, ICML, and ISMB, and the platform courses on YouTube. Placeholder narrative pending the owner's personal write-up.

Conductor & Composer

Conductor & Composer

This card is reserved for a closer look at contributions to the Bioconductor ecosystem and cloud-orchestrated bioinformatics workflows. Placeholder narrative pending the owner's personal write-up.

Writer placeholder

Writer — Technical & Prose

This card is reserved for a closer look at writing that sits outside pure documentation - the essays on Medium and posts on Gusmao Lab that mix technical and personal prose. See also Writing & Communication. Placeholder narrative pending the owner's personal write-up.


Something about the way chromatin folds always made me think about software. A single strand, roughly two meters of it, curls itself around histones, loops into topologically associated domains, and somehow keeps a coherent shape while every cell division unwinds it and rebuilds it, over and over, without corrupting the original state. It is not organized top-down. It emerges, correct.

I would like to know more...

I have spent a long time studying two kinds of structures that grow rather than get assembled: chromatin, and the platforms I build. Both start from something small and rule-bound - a sequence, a schema - and both end up more complex than any single rule predicted, if the underlying constraints were sound. Trees do this too. So do good codebases, on the rare occasions we let them.

A tree's roots are unglamorous and mostly invisible, and that is exactly why they matter. Get the architecture wrong - the data model, the ingestion contracts, the boundaries between systems - and no amount of visible polish above ground will hold. I have watched pipelines that looked finished on a dashboard collapse the first time an upstream schema changed quietly, because nobody had thought about the roots. A platform's roots are its architecture: not the diagram, but the actual set of guarantees the rest of the system is allowed to assume.

Branches are abstractions, and abstractions are a claim about the future - a bet that certain kinds of change will be common and others rare. Good branches grow toward light without breaking under their own weight; good abstractions absorb the variation a system actually encounters without becoming a museum of speculative flexibility nobody uses. I try to grow only the branches a platform's real weather pattern justifies, and prune the rest before they calcify into debt.

Leaves are features - individually small, replaceable, numerous, and that is fine, because that is what leaves are for. No single leaf makes a tree; no single feature makes a platform. What matters is whether the tree keeps producing them, season after season, without straining the structure that holds it up.

Bloom, when it happens, is not a separate event bolted onto the tree. It is what roots, branches, and leaves look like once they have had enough seasons to mature together - the moment a platform stops needing its author in the room to keep functioning, when the same rigor and reproducibility hold whether or not anyone is watching. That is the only kind of engineering maturity I actually trust: not a launch date, but a system that has had time to become itself.

There is a version of this that is purely biological, too. I named a chromatin-analysis tool Bloom for a literal reason - because folding, for DNA as for trees, is how a few meters of raw sequence become something legible enough to regulate a cell, or a genome, or an entire organism's development. The pattern is the same one degree removed: order, at scale, emerging from constraints simple enough to state and rich enough to surprise you. That is the intersection I keep circling back to, whether the raw material is contribution history, transcription-factor binding sites, or four bars of a chord progression that resolve in a way you did not quite predict but immediately trust. Systems allowed to grow honestly tend to end up, eventually, somewhere worth arriving.


🚀 "If you ever change your mind about leaving it all behind, remember. Remember. No Geography." 🚀

Designed & Built - Eduardo Gusmao - 2025

Popular repositories Loading

  1. Olympus Olympus Public

    A unified framework for discovering, analyzing, integrating, and visualizing regulatory motifs and transcription factor binding sites across bulk, single-cell, and long-read sequencing modalities.

    Python 11 4

  2. Bloom Bloom Public

    A Framework for Chromatin Architecture Data Processing, Handling and Analysis

    Python 11 6

  3. Blacksmith Blacksmith Public

    A high-performance genotype analysis framework for streamlined quality control, variant graph construction, and interactive network visualization

    Python 10 5

  4. Wildlife Wildlife Public

    A unified deep learning framework for high-performance multimodal data imputation, integrating neural operators for tabular, EHR, imaging, audio, video, and biological datasets

    Python 10 5

  5. Musique Musique Public

    A unified transcriptomics analysis framework supporting bulk, single-cell, long-read, short-read, and spatial expression workflows with integrated normalization, quantification, modeling, and visua…

    Python 10 5

  6. Fabric Fabric Public

    A collection of Health Informatics algorithms and tools.

    C 10 5