Forest landscape background

Maya Venkatraman

Research Software Engineer at the AlQuraishi LaboratoryPart-time M.A. in Statistics, Columbia University (2025 - present)

About Me

I'm Maya Venkatraman, a Research Software Engineer at the AlQuraishi Laboratory at Columbia University.

Currently, I work on developing a billion-parameter genome language model trained on prokaryotic DNA. My work spans the full ML research stack—from data curation and infrastructure to model architecture design, training, and biological benchmarking. I am particularly interested in the challenges of modeling biological sequences at single-nucleotide resolution and ultra-long context lengths, which push the boundaries of current transformer architectures.

In parallel to my research, I am pursuing a part-time Master's in Statistics at Columbia University, supported by the departmental MA2PhD Fellowship, which identifies students likely to pursue doctoral study in computational fields. I believe that rigorous statistical training is essential for advancing interpretable, efficient, and principled models in computational biology.

My interests within machine learning are both deep and broad, and I am considering multiple areas of focus for my PhD. I am particularly excited about the potential of diffusion models for protein design, as well as the use of reinforcement learning to guide exploration of conformational space. I am also drawn to novel AI techniques such as Hierarchical Dynamic Chunking from the Gu lab, valuing the idea of end-to-end, jointly optimized models that require less heuristic intervention. I believe that approaches like these will be especially relevant in biology, where data is inherently hierarchical, noisy, and context-dependent. I also see interpretability as a fascinating frontier in biological modeling, potentially enabling researchers to reverse engineer molecular contacts or mechanisms from patterns in model attention.

In the past, I worked at Google Research, applying computer vision to dermatologic image classification (Derm on Lens). I also contributed on a 20% basis to genomics projects, including DeepVariant and DeepNull. Before that, I worked at YouTube Trust and Safety, building infrastructure for detecting abusive user behavior.

Research Interests

Machine Learning

  • Guidance for diffusion and flow matching to improve biological design
  • Mechanistic interpretability — can we probe LLMs to reverse engineer biological mechanisms?
  • Training methods that enable us to train at biology-scale context lengths, with single-nucleotide resolution — think Hierarchical Dynamic Chunking
  • Reinforcement learning — can we explore conformational space while disincentivizing aphysical behavior?

Biology

  • Developing serum diagnostics through clever statistics, ML and high throughput technologies
  • Real-time integration of experimental results into ML methods to accelerate discovery
  • Causal inference to uncover mechanisms in chronic disease
  • Robotic or cloud laboratories

Education

Master of Arts in Statistics

Columbia University (2025 - present)
Supported by the Departmental MA2PhD 33k Merit Scholarship

Relevant Coursework:

  • COMS 4771: Machine Learning
  • MATH 2500: Analysis and Optimization
  • STCS 6701: Probabilistic Models and Machine Learning

Bachelor of Science in Computer Science

Columbia Engineering — Salutatorian, Class of 2022, GPA 4.12

Relevant Coursework:

CS: Artificial Intelligence, Applied Deep Learning, Computer Science Theory, Advanced Programming
Math: Probability and Statistics, Linear Algebra, Multivariable Calculus, Proofs in Analysis
Bio: Cellular and Molecular Biology I & II, General Chemistry I & II

Research

Research Software Engineer at the AlQuraishi Laboratory

Core member of team developing a billion-parameter prokaryotic genome language model (GLM) by scaling a BERT-style transformer with an MLM objective. Explored selective learning and distributed training methods to enhance model training efficiency. Designed and implemented novel biological benchmarks and reformulated existing benchmarks to assess scaling laws in our GLM. Co-developed a novel architecture for co-generating protein structure and sequence using diffusion.

Research Advisor: Mohammed AlQuraishi, Assistant Professor of Systems Biology & Computer Science

Department Profile: Columbia Systems Biology

DiffusionDeep LearningStructural BiologyComputational BiologyLanguage Models

Industry

Google Research - Health AI

L4 SWE

2023 - 2024

Applied computer vision techniques to classify skin lesions through the "Derm on Lens" project. Led launch of the new model, supporting first-ever ophthalmologic image classification. Contributed to research at the intersection of deep learning and genomics, including DeepVariant and DeepNull.

Research Advisors: Farhad Hormozdiari and Kishwar Shafin

Featured in U.S. Dermatology Partners: "I Used Google Lens To Check for Skin Cancer—Here's What Happened"

Project releases: DeepVariant Release Notes

Computer VisionHealthcare AIGenomicsDeep LearningPython

YouTube Trust and Safety - ML Infra

L3 → L4 SWE

2021 - 2023

Built ML infrastructure to detect abusive user behavior, improving platform safety at scale.

ML InfrastructureSpam DetectionLow-latency SystemsScalable SystemsC++

Publications

In a past life, I worked in a chemical engineering lab, developing computational models (a mix of algorithms and optimization) to demonstrate how hydrogen electrolysis can be performed at minimized cost.

Advisor: Daniel Esposito, Associate Professor of Chemical Engineering

Lab: Solar Fuels Engineering Laboratory

In Progress

Impact of a consumer-facing, AI-powered informational tool on retrospective evaluation of skin concerns

Sayres, R., Jain, A., Venkatraman, M., et al.

Expected 2025

News

DeepVariant v1.9.0 achieves ~20% runtime reduction and improved accuracy with the new HG002-T2T truth set. The latest release includes faster inference through optimized tensor handling and updated training schemes; I contributed through exploration of novel model architectures, hyperparameter tuning, and training dataset compositions.

Project details: DeepVariant Release Notes

Recognized as a National Merit Finalist for receiving a perfect score on the PSAT and selected as one of six Newton students to win the scholarship, chosen from over 15,000 finalists nationwide.

Featured in Newton Patch News

Awards

Fellowships and Grants

  • 🎓NSF CSGrad4US Fellowship (2025)
  • 🏫Columbia GSAS MA2PhD 33k Merit Scholarship (2025)
  • 💰Bonomi Scholar - $5K Research Grant (2021)
  • 🌍Earth Institute Collaborative Research Grant - $5K (2021)

Industry Awards

  • 📝Coming Soon...

Get in Touch

Connect with me on LinkedIn or email me at maya.venkatraman1@gmail.com