ML & data systems · Decision science · Product ownership

Ananya Chembai

I’m a Data Engineer at PNC working across production data systems, applied ML, automation, and stakeholder-facing delivery. I take messy, high-stakes, or ambiguous problems from requirements through release. My work spans ML and AI, data infrastructure, decision science, policy analysis, and technical product ownership. I’m strongest where technical work meets real decisions: framing the problem, evaluating trade-offs, translating analysis into recommendations, and building solutions people can actually use.

Experience

Where I’ve worked

PNC Bank

Pittsburgh, PA

Data EngineerFeb 2026 - Present

Data Engineer AssociateMay 2025 - Feb 2026

  • Promoted within nine months after receiving an Exceeds Expectations performance rating; one of three engineers promoted from a 28-person team.
  • Owned the two-month resolution of a monthly-statement delivery failure caused by 12–13-hour Oracle materialized-view refreshes, from solution design through production release. Redesigned the data architecture around daily snapshot tables, cutting refresh time by ~75% to 3–3.5 hours.
  • Engineered an agentic validation workflow for an 800GB broker-dealer trading database, automating pre- and post-refresh quality controls and reducing hands-on validation time by ~88% per monthly refresh.
  • Delivered a Canadian-entity reporting system for two new AML sanctions-screening requirements, translating regulatory compliance needs into SQL logic, escalation-ready reporting, and production workflows for downstream review.
  • Owned delivery of a T+1 allocations report, working directly with clients to translate reporting needs into a single output for allocated and unallocated trades. The solution eliminated manual reconciliation and enabled timely reporting under T+1 settlement deadlines.

UBS

Hyderabad, India

Financial Analyst InternApr 2024 - Jul 2024

  • Migrated month-to-date P&L reporting from legacy SQL to PySpark on Databricks, accelerating query execution and reducing overall report-generation time by ~67%.
  • Modernized legacy dataset delivery by moving releases to Azure and implementing Git-based CI/CD pipelines, creating a repeatable, version-controlled deployment process.

Education

What I studied

May 2025

Carnegie Mellon University

Master of Science in Data Analytics for Science · QPA: 4.17/4.33

Relevant coursework

  • Computational Modeling, Statistics & Machine Learning
  • Large-Scale Computing for Data Science
  • Neural Networks & Deep Learning in Science
  • Data Analytics for Decision Making

Jul 2024

Vellore Institute of Technology

Bachelor of Technology in Computer Science and Engineering · CGPA: 8.88/10

Relevant coursework

  • Artificial Intelligence
  • Natural Language Processing
  • Parallel & Distributed Computing
  • Data Structures & Algorithms

Projects

Things I’ve built

Predictive Optimization of Reddit BigQuery Workloads

Partnered directly with Reddit through a CMU capstone to analyze 4.1M BigQuery jobs using PySpark. Uncovered zero-output workloads consuming 13.8% of total slot time and multi-input jobs operating at over 90% lower efficiency. Built regression models to prioritize opportunities, then presented recommendations for batching metadata jobs, rescheduling heavy workloads, and right-sizing slot reservations directly to stakeholders.

  • GCP
  • BigQuery
  • PySpark
  • Spark SQL
  • Regression
  • Anomaly Analysis
  • Workload Segmentation

MatchMadeInMed — Synthetic Clinical-Trial Scenario Explorer

Built a rare-disease clinical-trial prototype that matched patient profiles to trial eligibility criteria. Used conditional GANs to generate synthetic control groups and XGBoost to predict placebo and treatment responses, achieving R² 0.96 on held-out data. Delivered an interactive Streamlit workflow for reviewing patient matches and comparing simulated placebo and treatment responses.

  • Python
  • Conditional GANs
  • XGBoost
  • Synthetic Controls
  • Streamlit
  • Model Evaluation
  • Synthetic Data

Retrieval-Augmented Research Assistant

Built a RAG system over academic papers using SentenceTransformers and ChromaDB, creating a question-based semantic index for evidence-grounded GPT-4o-mini responses. Added an “IDK” abstention rule for questions unsupported by the retrieved context, then evaluated held-out answers with BERTScore, achieving 0.84 precision, 0.92 recall, and 0.87 F1.

  • RAG
  • SentenceTransformers
  • ChromaDB
  • GPT-4o-mini
  • BERTScore
  • LLM Evaluation

Fashion Intelligence Engine

Created a fashion intelligence platform for Dango Closet, an early-stage fashion dropshipping venture. Designed as a cross-platform data aggregator, the system scores fashion trends and ranks colors, patterns, and styles for new apparel designs. Architected and validated the full-stack system, directing implementation through agentic coding tools and structured prompts. Integrated an eight-source ingestion pipeline with trend scoring, knowledge graphs, and automated API and workflow tests.

  • Product Development
  • Decision Support
  • Workflow Design
  • Agentic Coding
  • API Integration
  • Knowledge Graphs
  • Automated Testing

Healthcare Cost Equity & Policy Simulation

Independently sourced and cleaned 22K+ AHRQ MEPS records. Identified key cost drivers through exploratory analysis and evaluated Random Forest, XGBoost, and LightGBM models for predicting out-of-pocket healthcare costs. Designed four targeted policy interventions and built Monte Carlo simulations to compare government cost, population reach, and reduction of out-of-pocket burden. Recommended an income-targeted policy as the strongest balance of equity and cost-efficiency, then presented the analysis, simulations, and recommendation to a 30+ person audience.

  • Python
  • Pandas
  • Random Forest
  • XGBoost
  • LightGBM
  • Monte Carlo
  • Policy Analytics
  • Data Visualization

Resume

Experience, in one page

Choose the version that best matches the role: ML/AI and data systems, or product and decision science.

Contact

Let’s connect

I’m always happy to talk about interesting problems, thoughtful products, and new opportunities.