Data Scientist · Analytics Engineer · ML Engineer

Transforming data into actionable insights through advanced analytics, machine learning, and AI solutions

Curriculum vitae

Full CV — 2026

Two pages · PDF · 604 KB

Download↓

Education

PhD, Genetics & Molecular Biology

UNICAMP · bioinformatics · 2018–2022

Certification

The Business of Health Care

PennMed · Wharton · four courses

Verify at Coursera

Overview

The through-line in one phrase

Where you started, where you are now, and the thread connecting them. This is the slide that tells them how to listen to the next eight.

1Four domains
  • Research — what it taught you
  • Applied AI — what it taught you
  • Marketing & banking — what it taught you
  • Healthcare — what it taught you
2Numbers worth stating
  • Headline metric
  • Headline metric
  • Headline metric
Nov 2016 — Nov 2017Research

Scientific Internship

Genetics & Molecular Analysis Lab (LAGM) · UNICAMP

Finding the sugarcane genes that appear exactly once — stable reference points in a genome that repeats itself constantly.

Skills developed
  • Linux
  • Excel
  • Scientific Writing
  • Basic Statistics
  • Data Visualization
  • Presentation
1The work
  • Led a research project identifying and characterizing single-copy genes in sugarcane.
  • Comparative analysis against rice, corn and sorghum as reference genomes.
  • Cross-referenced sugarcane transcriptome data to isolate exclusive, low-polymorphism genes.
  • Functionally annotated the genes to uncover their biological roles.
2How I worked
  • Built and ran the project workflows in a Linux environment.
  • Bioinformatics tools plus Perl scripts for processing.
  • Excel for data analysis and management.
3Outcome
  • Delivered insights that advanced sugarcane genomics.
  • Applications in agricultural research and development.
Mar 2018 — Dec 2020Research

Contaminant Detection in Genotyping Data

Doctorate, phase I · LAGM, UNICAMP

Catching every contaminated sample in genotyping data — then handing breeders a tool that runs the check themselves.

Skills developed
  • Linux
  • R
  • Statistics
  • PCA
  • Machine Learning
  • Simulation
  • Scientific Writing
  • Presentation / Teaching
  • Software Development
1The work
  • Developed an innovative analysis method to detect contaminants in genotyping data from biparental populations.
  • Combined genetic metrics, principal component analysis and clustering into a single detection procedure.
  • Reached 100% contaminant detection.
2How I worked
  • Rigorously validated against simulated populations with known contamination.
  • Executed in a Linux environment using specific bioinformatics tools and R / Python.
3Outcome
  • Shipped as polyCID, a user-friendly R Shiny tool for routine quality control in plant breeding programs.
  • Published in Frontiers in Plant Science, Oct 2021 — 10.3389/fpls.2021.737919
4Figures
Dec 2020 — Dec 2022Research

Genomic Prediction & Gene Networks

Doctorate, phase II · LAGM, UNICAMP

Predicting how much a grass will yield from its genome with machine learning — then exploring the discovered associated genes.

Skills developed
  • Linux
  • R
  • Python
  • Statistics
  • Machine Learning
  • Complex Networks
  • Scientific Writing
  • Presentation / Teaching
1The work
  • Ran an extensive assessment of methods for predicting growth and biomass production in Urochloa ruziziensis, benchmarking conventional statistics against machine learning.
  • Predicted five agronomic traits — green matter, dry matter, stem and leaf dry matter, and regrowth — across 33 clippings spanning wet and dry seasons.
  • Scored every model on Pearson correlation and MSE under 5-fold cross-validation repeated 100 times.
  • Went past prediction into the biology: identified the genes associated with the selected markers and modelled how they interact as a co-expression network.
2How I worked
  • Phenotyped 1,000 plants across 50 half-sibling families, then bulked DNA per family so a thousand samples became fifty sequencing libraries.
  • Stacked three feature-selection algorithms — gradient tree boosting, random forest and extremely randomized trees — and kept only the markers that two or all three agreed on.
  • That intersection cut the variables to 1% of the original while holding Pearson correlation around 0.9 and MSE at 0.013: the same accuracy from a hundredth of the data.
  • Built and ran the whole pipeline in a Linux environment using R and Python.
3Outcome
  • Ranked the surviving markers by Gini importance and mapped the genes sitting within 5 kb of each one against a transcriptome assembled de novo from 11 genotypes.
  • Built a co-expression network around those genes and read out their neighbours, centrality and GO enrichment — treating the traits as interacting systems rather than independent effects.
  • Contributed strategies for maximizing genetic gain while cutting the time and cost of a breeding cycle: a hundredth of the markers means far cheaper genotyping per plant.
  • Published in Frontiers in Plant Science, Dec 2023 — 10.3389/fpls.2023.1303417
  • 2nd place, Bioinformatics oral presentation — GBMeeting 2022, UNICAMP.
4Figures
Dec 2022 — Nov 2023Applied AI

AI Lawyer Assistant & Invoice Extraction

Freelance · Campinas, Brazil

Two production systems — one that made Brazilian jurisprudence searchable by meaning, one that read a hundred invoices a day and handed back six hours of it.

Skills developed
  • Python
  • Web Scraping
  • Airflow
  • LangChain
  • LLMs
  • Vector Databases
  • FastAPI
  • Docker
  • Google Cloud
  • System Integration
1AI Lawyer Assistant
  • Scraped judicial decisions from Brazil's most important tribunals using Python with BeautifulSoup and Selenium.
  • Ran the collection as scheduled Airflow routines with error-message handling, so a scraper that broke announced itself instead of quietly returning nothing.
  • Built the search engine on a ClickHouse vector database, embedding each decision's official summary with OpenAI ada-002 — so a lawyer could search by meaning rather than by keyword.
  • Summarised the jurisprudence with GPT-3.5 through LangChain, turning long rulings into something triageable at a glance.
2Invoice Data Extraction
  • Automated invoice processing for a textile company receiving 100+ invoices daily, removing manual data extraction and standardising inconsistent formats into one.
  • Summarised each PDF invoice's context with GPT-3.5 through LangChain, then processed and prepared the extracted fields for downstream use.
  • Reached 98% accuracy and cut six hours of manual work per day.
  • Served inference from a FastAPI endpoint, containerised with Docker and deployed to Google Cloud Run.
  • Integrated the system into the tooling the company already ran.
3How I shipped it
  • Both systems ran unattended — one on a schedule with failure alerting, the other as an always-on API behind a container.
  • Treated the language model as one component in a pipeline rather than the product: scraping, orchestration, embedding, storage and integration were the bulk of the work.
  • Built in 2023, when GPT-3.5 and LangChain were brand new — the very start of the generative AI boom.
May 2024 — Sep 2024Marketing

CRM Campaign Data Validation

Data Scientist · Marketdata — VISA / Itaú

Making sure the numbers were right — CRM campaign data reconciled across three systems that all had to tell the same story.

Skills developed
  • SQL
  • AWS Athena
  • Salesforce
  • QuickSight
  • Excel
  • Data Quality
  • Funnel Analysis
  • Report Writing
1The work
  • Validated CRM campaign data for Itaú across Salesforce, AWS databases and QuickSight — the system of record, the SQL layer beneath it and the BI layer on top all had to agree.
  • Authored detailed reports on data quality and integrity, documenting what was wrong and where.
  • Extracted, processed and summarised campaign results using AWS Athena (SQL) and Excel.
  • Built dashboards and ran campaign funnel analysis, tracking where contacts fell out between stages.
2How I worked
  • Checked the same question at three levels — what Salesforce recorded, what the underlying AWS tables held, and what QuickSight displayed — since a mismatch at any one of them changes the conclusion.
  • Wrote SQL in Athena to pull campaign results directly, then carried the analysis into Excel, where the campaign team already worked.
  • Turned the checks into written reports on data quality rather than raw lists of discrepancies.
3Outcome
  • Campaign teams got numbers they could act on, with the caveats written down instead of assumed.
  • Took every discrepancy back to the team that owned it, showing where it originated and agreeing the fix — so corrections landed at the source instead of being patched downstream.
Sep 2024 — Nov 2025Marketing

Marketing Mix Modeling

Data Scientist · Marketdata — Gain Theory / Nomad Foods

Working out what the advertising was actually worth — regression models attributing frozen-food sales to media spend, so budget decisions had evidence under them.

Skills developed
  • Python
  • Excel
  • Statistics
  • Machine Learning
  • Marketing Mix Modeling
  • Data Visualization
  • Agentic Coding
  • English
1The work
  • Delivered Marketing Mix Modeling for Nomad Foods as part of the Gain Theory consultancy team, across the UK, Germany, Italy, France, Austria and Serbia.
  • Modelled a frozen-food portfolio — chicken and fish bites, vegetables and the rest of the range — against media spend on television, radio, YouTube, social media and out-of-home.
  • Fitted multiple linear regression, both OLS and Bayesian, using the consultancy's proprietary modelling tools.
  • Calculated marketing ROI per channel to evaluate campaign effectiveness.
  • Validated and prepared the input data — a mix model is only ever as good as the spend and sales series feeding it.
2How I worked
  • Worked in English with a distributed consultancy team, aligning objectives and deliverables across every market.
  • Improved the recurring workflow in Python and Excel, making each modelling cycle faster and more reliable rather than repeating it by hand.
  • Summarised findings as visualisations and presentations an international team could act on.
3Outcome
  • Per-channel ROI became the basis for how budget was split across television, digital and out-of-home.
  • Workflow improvements turned a manual, repeated process into a faster and more dependable one.
Nov 2025 — Jul 2026Healthcare

Medical Staffing Dimensioning

Data Scientist · Hapvida — São Paulo, Brazil

Replacing a spreadsheet with a system — hourly demand forecast and the roster solved for 144 hospitals and emergency units, meeting service levels without paying for idle hours.

Skills developed
  • Databricks
  • Linux
  • PySpark
  • SQL
  • Python
  • Machine Learning
  • Statistics
  • MLflow
  • Operational Research
  • Simulation
  • Streamlit
  • Software Engineering
  • Agentic Coding
1The work
  • Built an automated medical staffing system on Databricks that replaced a manual, spreadsheet-based process across 144 hospitals and emergency care units.
  • Produced data-backed, auditable rosters — every number traceable to the data that produced it, rather than to whoever last edited the sheet.
  • Targeted the real trade-off directly: meet the service level agreements while cutting idle medical hours, two objectives that pull against each other.
2Modelling
  • Forecast hourly demand with LightGBM, tuned using Optuna and backtested with walk-forward cross-validation — the only honest way to score a time series, since a random split lets the model see its own future.
  • Solved the roster as a mixed-integer linear program with PuLP, turning the forecast into an actual schedule under real staffing constraints.
  • Verified the solution with Monte Carlo simulation, so the roster was tested against demand that did not go to plan.
3Platform & engineering
  • Extracted from Oracle, engineered features and curated Delta tables in Unity Catalog, with data quality alerts and daily monitoring of system results against actuals, in SQL and PySpark.
  • Orchestrated and deployed as code with Databricks Asset Bundles, with experiment tracking in MLflow.
  • Shipped a Databricks BI dashboard so the business could self-serve outputs, history and KPIs, plus a Streamlit app to simulate scenarios and quantify what a given SLA target actually costs.
  • Treated it as software, not a notebook: documentation site with MkDocs, unit tests with pytest, code quality via Ruff, MyPy and pre-commit, and AI agent instructions and skills committed to the repository.
4Outcome
  • A manual process for 144 units became a monitored pipeline that runs on its own and reports when it disagrees with reality.
  • The business stopped asking for the numbers and started reading them, through the dashboard and the scenario app.
  • Staffing decisions became defensible — the roster can be explained, reproduced and audited after the fact.
5Figures
Jul 2026 — nowHealthcare

Delinquency Collection Scoring

Data Scientist · Hapvida — São Paulo, Brazil

Scoring overdue health-plan contracts by their probability of payment over time — so the collection team can build data-driven strategies instead of working the list by intuition.

Skills developed
  • Survival Analysis
  • Python
  • Statistics
  • Databricks
  • Linux
  • SQL
  • PySpark
  • Machine Learning
1The work
  • Developing a survival model that scores overdue health-plan contracts by their probability of payment over time.
  • Supporting the business team in targeting its collection strategy, to prevent cancellation by non-payment before it happens.
2Why survival analysis
  • A classifier answers will they pay. Collections needs will they pay, and when — a contract likely to settle next week and one likely to settle in two months call for entirely different action.
  • Survival models handle contracts that have not resolved yet without discarding them, which is most of the live portfolio at any moment.
3Where it is going
  • In progress — current work as of this deck.
  • The measures that matter are the business ones, not model fit: reduction in median delinquency time, and how much of the value at risk is recovered rather than lost to cancellation.
Self-directed

Portfolio

Nine projects, every one public on GitHub — two of them running as live apps.

Applications

2024/25 · LLM · Python · Streamlit
MAIA — My Artificial Intelligence App

Document chat, summarisation, image generation, object detection and agentic NL2VIZ workflows, on GPT-4, DALL·E and YOLOv8.

2024 · Python · FastAPI · Cloud Run
MLCenter

Predictive analytics hub serving the two models below as live endpoints.

Machine Learning

2022 · ML · Python
Fraudulent Transaction Detection

Precision 0.96, recall 0.73 — $57M revenue across 52,185 transactions.

2023 · ML · Python
Credit Risk Modeling

Precision 0.96, recall 0.74 — $38M portfolio at 85% acceptance, 8.6% bad rate.

2024 · ML · Python · Power BI
Energy Production Forecasting

ARIMA quarterly forecasts on ONS and Eneva production data, SMAPE 0.488.

2024 · ML · Python
Customer Segmentation

K-means on order frequency, timing and volume — five clusters built for targeting.

Analytics & SQL

2023 · SQL Server · Excel
Pizzeria Sales Dashboard

Revenue, average order value, daily and hourly trends, best and worst sellers.

2023 · SQL
Digital Platform Sales Analysis

Sales analysis for a digital-products marketplace, written up as a decision report.

2022 · SQL
Optimizing Online Sports Retail Revenue

CTEs and correlation across pricing, reviews and traffic to find revenue levers.

Explaining the work

Talks & Teaching

Conference talks and university lectures — every one with slides or video attached.

Thesis & research

2023 · thesis defenceEN
PhD Defense

Contaminant identification and multiomic analysis applied in polyploid tropical forage grasses molecular breeding.

2021 · conferenceEN
Contaminant Identification in Polyploids

Video-poster for the 11th Brazilian Congress of Plant Breeding, and the full talk at the IV GBMeeting — second place in the Bioinformatics section.

Genomics lectures

2022 · BV782, UNICAMPEN
Bioinformatics and AI Applied to Genomics

History, importance and future of bioinformatics, then an introduction to machine learning — supervised and unsupervised, with their evaluation metrics and plant-genetics examples.

2020 · BV782, UNICAMPPT-BR
Genomic Selection

From Meuwissen (2001) to how genomic selection models are evaluated, closing on my own work in half-sibling families of Urochloa ruziziensis.

Breeding & statistics

2019 · BV880, UNICAMPPT-BR
Genetic Markers

Morphological and molecular markers, focused on microsatellites and SNPs and their use in linkage maps and QTL mapping.

lab sessionEN
Linear Regression

Written as a scheme for my own study, then used to teach lab colleagues: fitting the line, calculating R², and the F-test.

Continuous learning

Certificates & Courses

42 certificates · 337+ documented hours · 2017 to 2026 — each title links to the certificate.

Machine Learning & MLOps5

2024Datacamp · 44h
2023Datacamp · 20h

Analytics, BI & Cloud4

2023Datacamp · 17h
2023Datacamp · 16h
2023Datacamp · 4h
2022Datacamp · 2h

SQL & Big Data9

2023Datacamp · 4h
2023Datacamp · 4h
2023Datacamp · 4h
2022Datacamp · 4h
2022Datacamp · 4h
2022Datacamp · 4h
2022Datacamp · 2h

Finance & Healthcare5

2026PennMed Wharton · 4 courses

Teaching — course assistant2

2021
Introduction to Python Programming
GenMelhor UFV · 16h

Statistics & Genomics10

2022GBMeeting UNICAMP · 20h
2021Conecta GEM · 6h
2021
Applied Quantitative Genetics
Illinois Urbana-Champaign
2021
Genomics-Assisted Breeding in Polyploids
USDA NIFA
2017UNIFESP · 22h

Johns Hopkins on Coursera3

2018Johns Hopkins
2018Johns Hopkins
What I reach for

Toolkit

Solid bars are roles, coloured by the kind of work; outlined bars are self-directed use. Ticks mark the years I hold a course or certificate — so where study ran ahead of use, you can see it.

20162017201820192020202120222023202420252026
Languages & data
Python
5.2y·
SQL
1.4y·
R
4.8y
Linux
6.9y·
Platforms & engineering
Databricks
1.1y·
PySpark
1.1y·
GitHub
7.0y·
Airflow
1.0y
Docker
1.0y
Agentic Coding
1.8y·
Modelling & AI
Statistics
7.0y·
Machine Learning
7.0y·
LLMs
1.0y
Simulation
3.4y·
Survival Analysis
0.5y·
ResearchApplied AIMarketingHealthcareself-directedcourse or certificate· still current
F. B. Martins