Knowledge Domains | Data Engineering & Platform Systems across cloud, hybrid, and regulated data environments
Technical Quality | Focused on production-grade pipelines: from raw ingestion -> trusted datasets -> analytics-ready products
Engineering Practice | Strong emphasis on reliability, reproducibility, auditability, and systems that age well
Broad Curriculum | Experience spanning ETL/ELT, distributed processing, cloud/HPC, data quality, and ML-adjacent systems
I would like to know more...
Hello, and welcome to my profile. My name is Eduardo - grab a cup of coffee and allow me to introduce myself.
I build data platforms and production-grade data systems where messy real-world inputs are transformed into reliable, auditable, and analysis-ready products.
In practical terms, my work lives somewhere between:
- Data ingestion - where reality arrives poorly formatted and with opinions.
- Transformation layers - where Python, SQL, Spark, and modeling discipline try to restore civilization.
- Data quality - because a pipeline that runs and a pipeline that is correct are not the same animal.
- Delivery - where analytics, BI, ML, and internal users need datasets that are stable enough to trust and boring enough to maintain.
I specialize in Python, SQL, PySpark/Spark, cloud and hybrid data infrastructure, ETL/ELT pipelines, data modeling, validation frameworks, and reproducible platform workflows. I have worked across environments involving AWS, GCP, Azure, HPC clusters, Docker/Kubernetes, CI/CD, and large heterogeneous datasets that rarely introduce themselves politely.
My engineering bias is simple: build systems that are clear, testable, observable, and maintainable after the original excitement has left the room.
This usually translates to:
- Production data pipelines with strong reproducibility, monitoring, and failure handling
- Scalable batch and distributed processing for high-volume analytical workloads
- Data modeling layers supporting analytics, BI, reporting, and ML workflows
- Validation and governance practices for environments where correctness is not decorative
- Cloud, hybrid, and HPC workflows designed to survive both scale and human memory
My background includes complex healthcare and enterprise data environments, but my current professional identity is straightforward: Senior Data Engineer / Data Platform Engineer. Machine Learning is still part of the toolbox, but the main job is now the plumbing, contracts, orchestration, and reliability that make downstream intelligence possible.
I value clean design, explicit trade-offs, and systems that are understandable by humans - not just machines with suspicious confidence.
Ethics, reproducibility, and long-term sustainability are not optional; they are part of the job.
Availability | Currently open to remote, hybrid, relocation-friendly, and long-term Data Engineering / Data Platform roles. See contact details. Relocation and onboarding take planning - good systems (and good moves) benefit from doing things properly.
2025 | Committer | Awarded Apache Spark Committer Status | The Apache Software Foundation (ASF) | Finland & Brazil
2022 | Senior Transition | Senior Data Engineer | Turku Biosciences & Brazilian Ministry of Health | Finland & Brazil
2020 | Outreach | Award-Winning COVID-19 Outreach Campaign | Göttingen General Hospital | Germany
2017 | Patent | LAG3-Targeting Cancer Therapy | Current owner: Bristol Myers Squibb | USA
2016 | Industry Transition | Data Engineer | Dana-Farber Cancer Institute | USA
2013 | Research | Computational Biology Researcher | RWTH University Medical School | Germany
I would like to know more...
|
|
name: "Eduardo Gusmao"
role: "Senior Data Engineer | Data Platform Engineer"
contact: "Recife, Brazil | eduardogade@gmail.com | github.com/eduardogade | linkedin.com/in/eduardogade"
languages: "English fluent | Portuguese native | Spanish B2 | German A2 | Finnish A1"
education: "2x PhD in Biomedical Informatics and Data Engineering / Computational Life Sciences; BSc + MSc in Computer Science"
summary: "Data Platform Engineer with 8+ years designing scalable data platforms, distributed systems, and production-grade pipelines across healthcare, life sciences, and enterprise environments. Strong Python, SQL, PySpark/Spark, AWS, Docker/Kubernetes, CI/CD, data quality, and analytics/BI platform experience."
professional_engagements:
current_role:
company: "Turku Biosciences / Brazilian Ministry of Health"
title: "Senior Data Engineer"
location: "Finland / Brazil"
date: "Sep 2022 - Present"
scope:
- "Lead a national-scale precision-medicine data platform integrating genomic, phenotypic, and clinical EHR data for 65,000+ individuals."
- "Build scalable Python, SQL, PySpark/Spark, Databricks-adjacent, HPC/SLURM, and API-driven ingestion and transformation workflows."
- "Deliver regulated ingestion, validation, governance, PII-compliant processing, observability, idempotency, and data quality controls."
- "Support analytics, BI, and ML workloads through reusable integration layers, backend data services, and optimized Parquet-based processing."
development_environment:
infrastructure: "AWS | Azure | GCP | HPC/SLURM | Docker | Kubernetes | Terraform | GitHub Actions"
languages: "Python | SQL | PySpark | Bash/Shell | Scala | Java | C/C++ | YAML | HCL"
data_stack: "PySpark | Spark | Pandas | Polars | NumPy | BigQuery | PostgreSQL | Parquet | JSON | dbt | dimensional modeling"
platform_engineering: "ETL/ELT | ingestion frameworks | transformation layers | platform APIs | CI/CD | observability | validation | data quality"
ml_ai_stack: "ML pipelines | MLOps | feature engineering | LLM APIs | embeddings | RAG | Hugging Face"
collaboration: "Agile/Scrum | stakeholder enablement | analytics teams | data scientists | engineers | product | infrastructure | security"I would like to know more...
name: "Eduardo Gusmao"
role: "Senior Data Engineer | Data Platform Engineer | Cloud Data Engineer"
location: "Recife, Brazil"
contact: "eduardogade@gmail.com | github.com/eduardogade | linkedin.com/in/eduardogade"
languages: "English fluent | Portuguese native | Spanish B2 | German A2 | Finnish A1"
summary: "Data Platform Engineer with 8+ years of experience designing and operating scalable data platforms, distributed data systems, and production-grade pipelines across healthcare, life sciences, and enterprise environments. Strong expertise in Python and SQL, with hands-on experience in PySpark/Spark, Kafka-style event-driven workflows, Airflow, AWS, Docker/Kubernetes, Terraform, and CI/CD to build reliable data infrastructure and platform services."
core_expertise:
- "Data Engineering"
- "Data Platform Engineering"
- "Cloud and Hybrid Data Infrastructure"
- "Distributed Data Systems"
- "ETL/ELT Pipelines"
- "Data Modeling and Analytics Engineering"
- "Data Quality and Governance"
- "Healthcare and Life Sciences Data"
- "Machine Learning Data Pipelines"
- "Production Reliability and Observability"
career_profile:
- "8+ years building scalable data platforms and distributed data systems"
- "Production-grade ingestion, transformation, validation, observability, and internal tooling"
- "Strong Python, SQL, PySpark/Spark, cloud, CI/CD, Docker/Kubernetes, and data quality background"
- "Experience supporting analytics, BI, ML workflows, and mission-critical data products"
- "Comfortable translating complex stakeholder requirements into maintainable platform capabilities"
professional_engagements:
current:
company: "Turku Biosciences / Brazilian Ministry of Health"
title: "Senior Data Engineer"
location: "Finland / Brazil"
date: "Sep 2022 - Present"
scope:
- "Lead the design and delivery of a national-scale data platform for precision medicine, integrating multi-modal genomic, phenotypic, and clinical EHR data for 65,000+ individuals."
- "Build scalable Python-based pipelines, distributed systems, PySpark/Spark workflows, Databricks-adjacent processing, HPC/SLURM execution, and API-driven ingestion workflows."
- "Architect production-grade data platform services for regulated ingestion, validation, transformation, governance, PII-compliant processing, reproducible workflows, and quality/reliability controls."
- "Develop high-performance processing and modeling layers using Python, SQL, PySpark, partitioning strategies, Parquet formats, and distributed query tuning."
- "Design reusable integration layers and backend data services connecting heterogeneous clinical, genomic, and ERP/SAP data sources."
- "Enable event-driven workflows, orchestration patterns, and batch/streaming-adjacent pipelines supporting analytics, BI, and ML systems."
- "Collaborate with product, analytics, engineering, and infrastructure stakeholders to deliver platform capabilities, CI/CD, Docker/Kubernetes workloads, observability, logging, alerting, and performance tuning."
outcomes:
- "Integrated precision-medicine datasets for 65,000+ individuals."
- "Improved pipeline efficiency by approximately 25%."
- "Reduced storage costs by approximately 80%."
- "Enabled more than 40% faster data delivery for downstream analytics, BI, and ML systems."
previous_mid:
company: "Göttingen General Hospital"
title: "Data Engineer II"
location: "Germany"
date: "Mar 2019 - Sep 2022"
scope:
- "Designed and implemented scalable data platform services using Python and SQL on cloud and hybrid environments."
- "Enabled reliable ingestion, transformation, and low-latency access for downstream analytics, BI, and application workloads."
- "Developed and optimized high-performance ETL/ELT pipelines with Python and PySpark, leveraging distributed processing, batch workflows, and orchestration patterns."
- "Refactored legacy systems into modular, production-grade platform services with CI/CD, automated testing, monitoring/logging, idempotency, retries, and robust error handling."
- "Built reusable data processing frameworks and integration layers for large-scale heterogeneous datasets."
- "Applied data modeling, validation, lifecycle standards, and governance across 12 cross-functional teams in a distributed environment."
outcomes:
- "Improved data availability and system responsiveness by approximately 33%."
- "Improved reliability, maintainability, and operational efficiency by approximately 50-60%."
- "Supported consistent data lifecycle practices across 12 cross-functional teams."
previous_old:
company: "Dana-Farber Cancer Institute"
title: "Data Engineer I"
location: "USA"
date: "Jan 2016 - Mar 2019"
scope:
- "Developed cloud-native data platform services supporting large-scale drug discovery."
- "Built Python-based ETL/ELT pipelines and API-driven integration layers for heterogeneous biomedical, operational, and financial datasets."
- "Implemented end-to-end data processing pipelines using Python, SQL, and PySpark on Apache Spark distributed systems."
- "Enabled scalable ingestion, transformation, validation, and batch workflows for analytics and ML-driven applications."
- "Collaborated with product, analytics, and research stakeholders to define KPIs and translate requirements into data models, backend data logic, and reusable platform components."
- "Contributed to production-grade data engineering practices including Git version control, validation checks, documentation, maintainable system design, reliability, and reproducibility."
outcomes:
- "Improved data accessibility and reduced operational costs by more than 25%."
- "Supported analytics and ML-driven applications through reusable data platform components."
- "Established reliable, reproducible lifecycle standards for heterogeneous biomedical and operational data."
education:
phd_biomedical_informatics:
degree: "Ph.D. in Biomedical Informatics"
institution: "Harvard Medical School"
location: "Boston / Cambridge, USA"
date: "2013 - 2017"
phd_computational_life_sciences:
degree: "Ph.D. Dr. rer. nat. in Data Engineering and Computational Life Sciences"
institution: "RWTH Aachen University"
location: "Aachen, Germany"
date: "2011 - 2015"
bachelor_master_computer_science:
degree: "B.Sc. and M.Sc. in Computer Science"
institution: "Federal University of Pernambuco"
location: "Recife, Brazil"
date: "2008 - 2011"
technical_strengths:
programming:
primary: ["Python", "SQL", "PySpark", "Spark SQL", "Bash/Shell"]
secondary: ["Scala", "Java", "C/C++", "YAML", "HCL"]
concepts: ["REST APIs", "Async programming", "Data serialization", "Production-grade software engineering", "Parquet", "JSON"]
data_platform_engineering:
capabilities: ["Scalable data platforms", "Distributed data systems", "Internal data tooling", "Reusable ingestion frameworks", "Transformation layers", "Platform APIs", "Developer-facing abstractions", "Self-service data capabilities", "Analytics enablement", "ML workflow support", "BI workload support"]
distributed_data_systems:
tools: ["PySpark", "Apache Spark", "Pandas", "Polars", "NumPy"]
capabilities: ["Large-scale processing", "Distributed compute", "Performance tuning", "Partitioning", "Query optimization", "Resource efficiency", "Batch pipelines", "Streaming-adjacent pipelines", "Kafka", "Spark Streaming patterns"]
cloud_hybrid_infrastructure:
cloud: ["AWS", "Azure", "GCP"]
infrastructure: ["HPC/SLURM", "Docker", "Kubernetes", "Terraform", "GitHub Actions", "GitLab CI"]
capabilities: ["Cloud-native data infrastructure", "Hybrid data infrastructure", "Infrastructure-aware engineering", "Containerized workloads", "Deployment environments", "Scalable platform operations"]
hadoop_on_prem_ecosystems:
technologies: ["HDFS", "YARN", "Hive", "Kerberos"]
capabilities: ["Distributed storage patterns", "Distributed compute patterns", "Legacy-to-modern platform evolution", "Secure access-controlled data environments"]
data_modeling_tooling:
: ["Dimensional modeling", "Semantic modeling", "Schema design", "Metadata management", "Transformation layers", "Data contracts", "Lineage", "Modeling standards", "Analytics enablement", "Platform consistency"]
tools: ["dbt", "BigQuery", "PostgreSQL", "Parquet", "JSON"]
software_engineering_devops_reliability:
tools: ["Git", "GitHub", "GitHub Actions", "GitLab CI", "Docker", "Kubernetes", "Terraform"]
practices: ["CI/CD pipelines", "Automated testing", "Deployment automation", "Monitoring", "Logging", "Alerting", "Observability", "Incident response", "Idempotency", "Retries", "SLA/SLO thinking", "Fault-tolerant design"]
machine_learning_data_pipelines:
capabilities: ["ML pipelines", "MLOps", "Feature engineering", "Data preparation", "Personalization workflows", "AI-enabled data workflows", "Production-oriented ML data support"]
ai_llm: ["LLM APIs", "Embedding pipelines", "RAG", "Hugging Face"]
data_security_governance_quality:
capabilities: ["Data privacy", "PII-aware processing", "Compliance-aware pipelines", "Access control", "Validation strategies", "Auditability", "Data quality checks", "Governance practices", "Secure data lifecycle management", "Reliability controls", "Consistency checks"]
processes_collaboration:
practices: ["High-ownership engineering mindset", "Agile/Scrum", "Cross-functional collaboration", "Stakeholder enablement", "Requirements translation", "Technical documentation", "Platform capability delivery"]
collaborators: ["Analysts", "Data scientists", "Engineers", "Product teams", "Infrastructure teams", "Security teams"]
development_environment:
hardware: ["Apple Silicon", "ARM", "Intel", "NVIDIA GPU environments", "HPC clusters"]
operating_systems: ["macOS", "Ubuntu", "Debian", "Fedora", "Windows"]
infrastructure:
cloud_computing: ["AWS", "Azure", "GCP"]
hpc: ["SLURM", "OpenPBS", "Distributed compute environments"]
containers: ["Docker", "Kubernetes", "Singularity"]
infrastructure_as_code: ["Terraform", "HCL", "Cloud deployment"]
languages:
data_engineering: ["Python", "SQL", "PySpark", "Spark SQL", "Bash/Shell"]
systems_and_general: ["C/C++", "Java", "Scala"]
markup_and_config: ["YAML", "Markdown", "LaTeX", "HTML/CSS", "HCL"]
data_stack:
distributed_processing: ["Apache Spark", "PySpark", "Spark SQL", "Pandas", "Polars", "NumPy"]
storage_formats: ["Parquet", "JSON", "CSV", "HDF5"]
databases_and_warehouses: ["BigQuery", "PostgreSQL", "MongoDB", "DynamoDB", "Relational databases", "NoSQL databases"]
modeling_and_quality: ["Dimensional modeling", "Semantic modeling", "Schema design", "Data contracts", "Lineage", "Validation checks", "Data quality checks", "dbt"]
ml_ai_stack:
frameworks_and_tools: ["PyTorch", "TensorFlow", "Keras", "Scikit-Learn", "Hugging Face", "NLTK"]
workflows: ["ML pipelines", "MLOps", "Feature engineering", "Embedding pipelines", "RAG", "LLM APIs"]
systems_tooling:
version_control: ["Git", "GitHub"]
packaging_and_environments: ["pip", "poetry", "micromamba", "mamba", "conda", "npm"]
ci_cd: ["GitHub Actions", "GitLab CI"]
observability: ["Logging", "Monitoring", "Alerting", "Observability", "Prometheus", "Grafana"]
github_positioning:
short_pitch: "I build reliable data platforms, distributed pipelines, and production-ready data systems for analytics, BI, ML, and healthcare/life-sciences workloads."
engineering_style:
- "Clean, maintainable, typed Python"
- "Data quality and reliability first"
- "Production-aware platform design"
- "Reproducible workflows"
- "Strong documentation"
- "Pragmatic cloud and hybrid infrastructure" LinkedIn | https://www.linkedin.com/in/eduardogade/
Location | Recife, Brazil | Remote-friendly
Status | Open to Data Engineering roles
I would like to know more...
LinkedIn: https://www.linkedin.com/in/eduardogade/
Website & Blog: https://www.gusmaolab.org
CV/Resume: https://www.gusmaolab.org/cv/CV_Eduardo_Gusmao.pdf
Stack Overflow: https://stackoverflow.com/users/32223943/eduardo-gusmao
Medium: https://medium.com/@eduardogade
Preferred contact: Email | LinkedIn
Response time: 1-2 business days
Open to remote, hybrid, or relocation
See [availability & engagement details](#availability)
|
Apollo A unified suite of post-hoc statistical procedures with bias-aware corrections designed for metrics common in computational and ML/DL pipelines. View Repository → |
Blacksmith A high-performance genotype analysis framework for streamlined quality control, variant graph construction, and interactive network visualization. View Repository → |
Olympus A unified framework for discovering, analyzing, integrating, and visualizing regulatory motifs and transcription factor binding sites across bulk, single-cell, and long-read sequencing modalities. View Repository → |
I would like to know more...
|
Apollo A unified suite of post-hoc statistical procedures with bias-aware corrections designed for metrics common in computational and ML/DL pipelines. |
Blacksmith A high-performance genotype analysis framework for streamlined quality control, variant graph construction, and interactive network visualization. |
Olympus A unified framework for discovering, analyzing, integrating, and visualizing regulatory motifs and transcription factor binding sites across bulk, single-cell, and long-read sequencing modalities. |
|
Bloom A framework for chromatin architecture data processing, handling and analysis. |
Musique A unified transcriptomics analysis framework supporting bulk, single-cell, long-read, short-read, and spatial expression workflows with integrated normalization, quantification, modeling, and visualization tools. |
Wildlife A unified deep learning framework for high-performance multimodal data imputation, integrating neural operators for tabular, EHR, imaging, audio, video, and biological datasets. |
|
Fabric A collection of health informatics algorithms and tools. |
Uqbar Ubiquitously Broad Automation and Architecture — a collection of tools for small task automation. |
GusmaoLab Source of Eduardo Gusmao's lab website — portfolio, technical blog, and public-facing documentation. |
Apache Software Foundation | Committer | Apache Spark | 2025 - Present
Maintainer | 9+ actively maintained research and engineering tools | See Pinned Repositories
Community & Volunteering | Global Burden of Disease, TransEmpregos, ABRATA | See Causes
Status | Curating additional public contribution history
I would like to know more...
Open-source involvement is treated as an extension of day-to-day engineering practice rather than a separate résumé line - it is where tooling, methods, and lessons from production work get generalized and given back.
The most concrete example is the Apache Spark committer status awarded by the Apache Software Foundation in 2025, which followed sustained engagement with the project's distributed-processing internals through professional and personal work. Beyond that, most of the ongoing contribution activity currently lives in the personal toolkit showcased under Pinned Repositories - actively maintained libraries for statistics, regulatory genomics, chromatin analysis, and small-scale automation - and in the volunteering commitments described under Causes.
A broader, evidence-backed view of pull requests, issue triage, and cross-project reviews across external organizations is still being organized for public presentation. Rather than publish an incomplete or padded picture, this subsection will be expanded once that material is ready to stand on its own.
I would like to know more...
Additional breakdowns of repositories, commit language distribution, and contribution timing - not already shown above.
This section is reserved for repository-influence and collaboration analytics - dependency relationships, external review activity, and community engagement across the tools in Pinned Repositories.
I would like to know more...
The dashboard above already covers language mix, commit activity, and contribution cadence well. What is intentionally not here yet is a second layer of analytics: how these tools get used and reviewed outside of direct authorship - dependency graphs, downstream adoption, and external review activity.
No currently available public tooling produces that view at a quality bar consistent with the rest of this page without either being self-hosted or relying on metrics (like early-stage star counts) that would be more decorative than informative for research tooling with a narrow, specialist audience. Rather than fill the space with a vanity widget, this section stays intentionally reserved until metrics that are actually meaningful - closer to genuine downstream engineering usage than surface-level GitHub counters - become available.
I would like to know more...
Complete technology inventory, synchronized with CV.yaml and grouped by domain. Icons shown where a stable public icon set provides one; unmarked names are listed as text only.
Programming Languages
Data Engineering & Orchestration
Databases & Storage
Cloud & Infrastructure
Machine Learning & AI
Scientific & Bioinformatics Computing
Web, APIs & Dashboards
Observability & DevOps
Documentation & Markup
Machine & Deep Learning | Repository | Publication
Variational Inference | Repository | Publication
Precision Medicine | Repository | Publication
Regulatory Genomics | Repository | Publication
I would like to know more...
Global age-sex-specific all-cause mortality and life expectancy estimates for 204 countries and territories and 660 subnational locations, 1950-2023: a demographic analysis for the Global Burden of Disease Study 2023
The Lancet · Oct 18, 2025
Contributions:
- Responsible for orchestrating the LATAM-branch with 45+ PIs and 200+ researchers.
- Horizontal meetings for data and experience sharing have shown great success, with ~380% more efficiency than the second most efficient branch - per capita.
- Has solved pharmacological conflict of interests by cross-deployment and blind-genotype blind-phenotype strategy, which exhibit 17% increased accuracy over North America (first COI - percapita) and 5% over Asia (second COI - per capita).
Cell Reports · Nov 26, 2024
Contributions:
- The tool
Bloomhas increased analysis mechanism by promoting different views into the regulatory spatial configuration, resulting in ~50% wet-lab equipment cost reduction and solving a stalled-case. - Provided personal guidance towards architecture and Hi-C methodology, saving 15% overall lab-time.
- Overall, this was the first non-trivial non-intermediary-distance (>1Gbp) lncRNA interference in a region unknown to be a regulatory enhancer.
Global, regional, and national burden of diabetes from 1990 to 2021, with projections of prevalence to 2050: a systematic analysis for the Global Burden of Disease Study 2021
The Lancet · Jul 15, 2023
Contributions:
- Responsible for orchestrating a team of 3 brazilian PIs and 5 independent investigators.
- Used scrum, coupled with CRISP-DM, delivering net gains (profitability converted back) through network revenue saving and wet/dry-lab material cost reduction.
- Developed national-scale geno/phenotype QC pipeline -
Fabric(PhenotekaModule) - used across 20+ institutes.
The New England Journal of Medicine · Oct 11, 2022
Contributions:
- Developed
Blacksmith, that coupled with Bloom improved operating margin by over 15%. - Freed at least 15 engineering hours per week with Blacksmith coupled with
Apollo. - Intending to lower carbon footprint, we adopted a trademarked DB 'bit-brushing' methodology (currently owned by Databricks Inc.).
Molecular Systems Biology · Jun 24, 2021
Contributions:
- Analized Spatial Chromatin Biology and RNA-seq to identify - for the first time - HMGB1 as a 'rheostat' factor.
- Reduced cloud compute costs by 40% using
Apollo's strong mathematical features and Bloom to analyse Chromatin conformation. - After this project's results, we have earned an ESG compliance through impeccable waste management and safety handling.
Redundant and specific roles of cohesin STAG subunits in chromatin looping and transcriptional control
Genome Research · Apr 6, 2020
Contributions:
- Analized most omics in a single project: ChIP-seq, degron-X, RNA-seq, Hi-C, STORM, DNase-seq, ATAC-seq, MSMS and MS-based microscopy.
- Developed Musique, shortening development cycles by ~9 weeks.
- Musique saved 300 GPU-hours per month by performing simple heuristics which are generalizeable to any dataset.
Spatial chromosome folding and active transcription drive DNA fragility and formation of oncogenic MLL translocations
Molecular Cell · Jul 25, 2019
Contributions:
- Patented technique for BLISS-seq data processing, earning ~25% extra funds for the laboratory.
- Lower wet-lab costs using dry-lab tools by ~30% (estimated for this project); achieving reproducible and insightful results on MLL fusions.
- Created the triple-correlation method. Translating category theory into a real-world phenomenon.
HMGB2 loss upon senescence entry disrupts genomic organisation and induces CTCF clustering across cell types
Molecular Cell · May 17, 2018
Contributions:
- Developed
BloomandApollo, which reduced processing time by at least 3 months. - Very agile methodology with microprocessed multicycled days, leading to novel discoveries and decreasing overall time-to-delivery.
- Reduced local infrastructure storage footprint by ~100TB with
Bloom&Apollo.
Nature · Jan 23, 2017
Contributions:
- Devised bioinformatics pipelines with collaborators and created the Gaussian-as-DPMM method of clustering, increasing speed by, at least, ~100x.
- Clustering was able to identify 3+ unique subtypes never previously reported.
- Created a deep regulatory network, especially with SHKBP1 ERBB3 and TGFBR2; which contained 98% of the cancer mortality information variability.
Nature Methods · Feb 22, 2016
Contributions:
- Landmark study on comparing 12+ footprinting methods. The study was the cover of Nature Methods magazine.
- Without any dry experiment, we were able to identify the limits of sequencing technologies, and propose results that exceded ~5% AUPR of known methods.
- Our method - Olympus (published in 2023) - offers ~7x most complete analysis of regulatory genomics than any other tool.
Nucleic Acids Research · Oct 17, 2015
Contributions:
- First use of
Faun, the motif enrichment analysis that uses hypergeometric distributions to query the sensitivity and specificity of TF occupancy in a certain genomic region. - Proposed the usage of
Cytoscape, widely minimising meeting preparation time by ~25%. - Proposed use of fewer histone modification essays by recreating chromatin states in silico; thus, minimizing project costs by ~30%
I would like to know more...
I treat writing as an engineering discipline, not an afterthought bolted onto the end of a project. A pipeline that nobody can explain six months later is not really finished, no matter how well it runs today - the explanation is part of the deliverable, and it degrades just like code does if nobody maintains it.
Most of what I write falls into two categories: documentation that has to survive contact with someone else's Monday morning, and architecture notes that have to survive contact with my own memory a year from now. Both audiences are unforgiving in the same way - they punish vagueness, reward precision, and do not care how elegant the underlying system is if the explanation doesn't hold up. Writing clearly about a system is often the fastest way to discover that it isn't as clearly designed as I thought.
The Atlas Learning Series and the posts on Gusmao Lab and Medium exist for the same reason internal design docs exist: complexity that stays in one person's head is a liability, and complexity that gets written down - honestly, without inflating it - becomes something a team can actually build on.
Status | Open to new opportunities | Data Engineering / Data Platform roles
Work style | Remote, hybrid, or on-site | Relocation-friendly
Geography | Europe → Brazil → Worldwide, roughly in that order of preference
Engagement | Long-term, technically demanding roles preferred over short-term contracts
I would like to know more...
Geography is a genuine preference, not a hard boundary. Europe and Brazil are where my professional and personal ties are strongest, so a role based there - or a remote/hybrid arrangement rooted there - tends to be the smoothest fit. That said, a worldwide, well-run, technically serious opportunity is always worth a conversation; relocation and onboarding just take a bit more planning, the same way any well-run migration does.
What I optimize for in a team is closer to craftsmanship than throughput. I would rather ship a system that is a little slower to build and considerably easier to operate, extend, and hand off, than one that is fast to demo and expensive to live with. That preference shows up as a bias toward explicit trade-offs, readable code, tests that actually catch regressions, and documentation that survives the person who wrote it moving on to something else.
Productivity, in my experience, tracks environment more than effort. Teams with reproducible builds, sane CI/CD, infrastructure-as-code, and automation around the boring parts of the job consistently ship better software than teams without them - independent of talent, tooling brand, or operating system. I look for that kind of environment because it is where careful engineering actually compounds instead of getting rebuilt from scratch every eighteen months, and I try to leave every team I join a little closer to it than I found it.
Collaboration matters just as much as any individual technical choice. The best systems I have worked on came out of teams that argued about trade-offs early, wrote decisions down, and treated architecture as something owned collectively rather than defended individually - that is the kind of team I look for, and the kind I try to build.
This section will eventually map out long-term technical direction rather than a short-term task list: platform evolution, research directions worth pursuing further, open-source initiatives planned for the tools under Pinned Repositories, and the engineering initiatives that connect them. The illustration above is a placeholder reserved for that map.
I would like to know more...
The expanded version of this section will lay out where the underlying platforms and research directions are headed over the next several years - which parts of the current toolkit graduate from personal research tools into broader shared infrastructure, which open questions from the publication record are worth a dedicated engineering push, and which engineering practices (observability, reproducibility, data contracts) are being generalized across projects rather than reinvented per repository. The illustration above will be replaced with the final roadmap artwork once that direction is settled.
2017 | PhD | Biomedical Informatics | Harvard Medical School
2015 | PhD | Life Sciences | RWTH Aachen University
2011 | MSc | Machine & Deep Learning | Federal University of Pernambuco
2008 | BSc | Computer Science | Federal University of Pernambuco
I would like to know more...
Ph.D. in Biomedical Informatics — Harvard Medical School (2013 - 2017)
Doctoral work carried out in parallel with a data engineering role at the Dana-Farber Cancer Institute, focused on computational pipelines supporting CRISPR-based immunotherapy target discovery. Contributed to target-discovery work that advanced into clinical development for LAG-3-based immunotherapy, later commercialized as Opdualag (Relatlimab + Nivolumab). Co-authored studies published during this period spanning cervical cancer genomics (Nature), chromatin organization in senescence (Molecular Cell), and computational footprinting methods for regulatory genomics (Nature Methods) - see Research.
Ph.D. Dr. rer. nat. in Data Engineering and Computational Life Sciences — RWTH Aachen University (2011 - 2015)
Doctoral research in regulatory genomics and epigenomics, centered on computational footprinting methods for DNase-seq and related chromatin assays. This period produced the analytical groundwork later published as a landmark comparison of 12+ footprinting methods (cover article, Nature Methods) and the method that became Olympus, alongside early chromatin-architecture work that would grow into Bloom.
M.Sc. in Computer Science (Machine & Deep Learning) — Federal University of Pernambuco (2010 - 2011)
This subsection will summarize the thesis focus, key coursework, and any early research output from the Master's program once that material is organized for public presentation.
B.Sc. in Computer Science — Federal University of Pernambuco (2008 - 2011)
This subsection will summarize notable coursework, early projects, and foundational technical milestones from the undergraduate program once that material is organized for public presentation.
Craftsmanship | Software is read far more often than it is written; optimize accordingly
Reproducibility | If a result cannot be rebuilt from scratch, it is not yet a result
Documentation & Testing | Explicit beats implicit; tested beats assumed
Data Governance | Especially with healthcare and other sensitive data, privacy and auditability are not optional
Scientific Integrity | Reproducible engineering practices are how research claims earn trust
I would like to know more...
Most of my engineering principles come from having worked on data that describes real people - patients, cohorts, research participants - long before it became rows in a table. That background makes certain things non-negotiable: sensitive data gets handled with explicit access controls and audit trails, not "we'll get to it later" promises, and a pipeline that cannot explain its own provenance is not trustworthy no matter how good its output looks.
Maintainability is the other constant. Code that only its author can safely change is a liability with a delay timer on it, and documentation that describes what a system does two versions ago is worse than no documentation at all, because it actively misleads. I would rather spend an extra afternoon on a clear interface and an honest README than leave that debt for whoever inherits the system - often enough, that person is me, eighteen months later, with no memory of the clever shortcut I took.
Scientific reproducibility and software engineering discipline turned out to be the same skill wearing different clothes. A result that cannot be rebuilt from raw inputs by someone else is not really a result yet, just a claim; the same containerized, version-controlled, tested rigor that makes a data pipeline trustworthy in production is what makes a research finding trustworthy in a paper. I try not to draw a hard line between the two.
A few thinkers show up in how I approach this, more as working tools than as heroes. Wittgenstein's attention to the limits of language maps directly onto naming things well and writing documentation that means exactly what it says - ambiguity in a spec is a bug, not a style choice. Camus's insistence on working honestly inside an absurd, uncertain world is a fair description of what building on messy real-world data actually feels like: you do not get certainty, you get discipline. And the habit of asking who actually benefits from a given system, and who quietly bears its costs, is a useful lens for thinking about the social impact of the technology I help build - not a political position, just a question worth asking before shipping.
Flagship: 🏳️⚧️ | 🏳️🌈 | 🇺🇳
|
Global Burden of Disease
Collaborator · 2018 - Present Contributing to one of the world's largest collaborative efforts for measuring disease burden through data infrastructure, engineering, and scientific collaboration. |
TransEmpregos
Volunteer · 2024 - Present Supporting infrastructure and data engineering for employability initiatives focused on transgender professionals. |
ABRATA
Volunteer · 2022 - 2025 Supporting initiatives that improve awareness and education regarding mood disorders and mental health. |
I would like to know more...
Global Burden of Disease is one of the largest ongoing scientific collaborations in public health, coordinating thousands of researchers across nearly every country to produce comparable estimates of disease, injury, and risk-factor burden over time. For a data engineer, it is also a genuinely hard distributed-collaboration problem - reconciling heterogeneous data sources, methods, and reporting standards across institutions that rarely share infrastructure. Participating in that network connects day-to-day pipeline work to public-health decisions made well beyond any single hospital or country.
TransEmpregos works to close the employment gap faced by transgender professionals in Brazil, one of the countries where that gap is widest. The technical need is unglamorous and familiar - reliable data infrastructure so the organization can run, match opportunities, and measure impact - but the outcome is concrete: real people getting a fairer shot at stable employment.
ABRATA (Associação Brasileira de Familiares, Amigos e Portadores de Transtornos Afetivos) works on public awareness and education around mood disorders in Brazil, where mental-health literacy still lags far behind physical-health literacy. Supporting an organization like this is a reminder that not every meaningful contribution has to be technical - sometimes it is simply showing up.
I would like to know more...
|
Open Source Contributor This card is reserved for a closer look at the personal research toolkit maintained under Pinned Repositories and the broader open-source habits behind it. Placeholder narrative pending the owner's personal write-up. |
|
Lecturer & Public Speaker This card is reserved for a closer look at teaching CS at UFPE/CIn and TUM/SLS, keynote talks at NeurIPS, ICML, and ISMB, and the platform courses on YouTube. Placeholder narrative pending the owner's personal write-up. |
|
Conductor & Composer This card is reserved for a closer look at contributions to the Bioconductor ecosystem and cloud-orchestrated bioinformatics workflows. Placeholder narrative pending the owner's personal write-up. |
|
Writer — Technical & Prose This card is reserved for a closer look at writing that sits outside pure documentation - the essays on Medium and posts on Gusmao Lab that mix technical and personal prose. See also Writing & Communication. Placeholder narrative pending the owner's personal write-up. |
Something about the way chromatin folds always made me think about software. A single strand, roughly two meters of it, curls itself around histones, loops into topologically associated domains, and somehow keeps a coherent shape while every cell division unwinds it and rebuilds it, over and over, without corrupting the original state. It is not organized top-down. It emerges, correct.
I would like to know more...
I have spent a long time studying two kinds of structures that grow rather than get assembled: chromatin, and the platforms I build. Both start from something small and rule-bound - a sequence, a schema - and both end up more complex than any single rule predicted, if the underlying constraints were sound. Trees do this too. So do good codebases, on the rare occasions we let them.
A tree's roots are unglamorous and mostly invisible, and that is exactly why they matter. Get the architecture wrong - the data model, the ingestion contracts, the boundaries between systems - and no amount of visible polish above ground will hold. I have watched pipelines that looked finished on a dashboard collapse the first time an upstream schema changed quietly, because nobody had thought about the roots. A platform's roots are its architecture: not the diagram, but the actual set of guarantees the rest of the system is allowed to assume.
Branches are abstractions, and abstractions are a claim about the future - a bet that certain kinds of change will be common and others rare. Good branches grow toward light without breaking under their own weight; good abstractions absorb the variation a system actually encounters without becoming a museum of speculative flexibility nobody uses. I try to grow only the branches a platform's real weather pattern justifies, and prune the rest before they calcify into debt.
Leaves are features - individually small, replaceable, numerous, and that is fine, because that is what leaves are for. No single leaf makes a tree; no single feature makes a platform. What matters is whether the tree keeps producing them, season after season, without straining the structure that holds it up.
Bloom, when it happens, is not a separate event bolted onto the tree. It is what roots, branches, and leaves look like once they have had enough seasons to mature together - the moment a platform stops needing its author in the room to keep functioning, when the same rigor and reproducibility hold whether or not anyone is watching. That is the only kind of engineering maturity I actually trust: not a launch date, but a system that has had time to become itself.
There is a version of this that is purely biological, too. I named a chromatin-analysis tool Bloom for a literal reason - because folding, for DNA as for trees, is how a few meters of raw sequence become something legible enough to regulate a cell, or a genome, or an entire organism's development. The pattern is the same one degree removed: order, at scale, emerging from constraints simple enough to state and rich enough to surprise you. That is the intersection I keep circling back to, whether the raw material is contribution history, transcription-factor binding sites, or four bars of a chord progression that resolve in a way you did not quite predict but immediately trust. Systems allowed to grow honestly tend to end up, eventually, somewhere worth arriving.








