Machine learning engineer

Master’s student in Computer Science, Rice University

I fine-tune and evaluate language and vision-language models, and build the infrastructure that keeps them reliable once they’re deployed.

See selected work

Selected work

Medical VLM reliability evaluation

An evaluation of Qwen2.5-VL-7B and Qwen3-VL-8B on a balanced set of 200 chest X-rays, zero-shot and fine-tuned with LoRA, scoring each diagnosis together with its stated confidence and explanation.

Accuracy and F1 hid the real problem, so we added metrics for overconfident errors, fluency and faithfulness. The strongest model still missed over half of pneumonia cases, and its wrong answers were its most fluent: 0.855 fluency on errors against 0.655 accuracy.

Role
Co-author: evaluation, metrics and fine-tuning
Stack
PyTorch · Hugging Face · Qwen-VL · LoRA · 4-bit NF4
Result
76.8% of Qwen3-VL’s errors made with high confidence
SourceView on GitHub (KaavinB/Medical-VLM-Reliability-Evaluation, opens in a new tab)

LLaMA fine-tuning for research classification

Fine-tuned LLaMA-3.2-3B to sort arXiv papers by research area from their title and abstract, using a balanced dataset of 2,000+ abstracts I curated.

Rank-16 LoRA adapters on a 4-bit NF4 base cut the trainable parameters to about 6M, so the whole thing trains on a local GPU. Zero-shot and adapted models run through the same inference pipeline for a fair comparison.

Role
Fine-tuning and data curation
Stack
PyTorch · Hugging Face · LoRA · 4-bit NF4
Result
40% → 67% classification accuracy
SourceView on GitHub (KaavinB/finetuning_arXiv, opens in a new tab)

MLOps sentiment pipeline on AWS

A sentiment classifier for 50K IMDB reviews, taken from raw data to a monitored endpoint on AWS EKS. DVC versions the data and MLflow tracks every run and holds the model registry.

One git push runs a 10-step GitHub Actions workflow: rebuild the pipeline, test the model and the server, promote the model to Production and roll the new image out to EKS through ECR. Prometheus and Grafana track latency, throughput and health, with p95 latency at 12.8 ms.

Role
Pipeline and infrastructure
Stack
Docker · AWS EKS · GitHub Actions · DVC · MLflow · Prometheus · Grafana
Result
71% → 87.7% test accuracy
SourceView on GitHub (KaavinB/mlops-sentiment-pipeline, opens in a new tab)

Experience

  1. May 2026 – Aug 2026

    Season 25/26

    AI Engineering Intern

    Rice Center for Engineering Leadership, Houston, TX

    Built a natural-language search tool that ranks 68 US AAU universities by research strength across 110K dissertations, for faculty hiring committees. Bayesian shrinkage corrected a small-sample bias that inflated specialization scores up to 18×, cutting unreliable rankings by 88%. Queries are matched to a 77-topic taxonomy with coarse-to-fine search over MiniLM embeddings and checked by an LLM, served through FastAPI on GCP Cloud Run.

  2. Aug 2025 – Dec 2025

    Season 25/26

    Teaching Assistant

    Rice University, Houston, TX

    Automata, Formal Languages & Computability

    Office hours and grading for a graduate automata theory course. Helping students work through formal proofs is harder to teach than it looks.

  3. Jan 2024 – Apr 2024

    Season 23/24

    Software Engineering Intern

    VIT Chennai, Chennai, India

    Built a face-recognition attendance system used by 500+ students on campus. Got inference under 0.5s and shipped a companion React Native app to production.

  4. Nov 2023 – Dec 2023

    Season 23/24

    Machine Learning Research Intern

    VIT Chennai, Chennai, India

    Built LSTM models for wind energy forecasting that beat the baseline by about 10%, and set up a federated training pipeline with Flower to keep raw data local.

Education

  1. Aug 2025 – Dec 2026 (expected)

    Rice University

    Master of Computer Science, Houston, TX

    GPA 3.68 / 4.0

  2. Aug 2021 – Apr 2025

    Vellore Institute of Technology

    B.Tech in Electronics and Computer Engineering, Chennai, India

    GPA 8.66 / 10

More work

Research

  1. AI-Powered Attendance System Using Facial Recognition

    IEEE Conference, 2024

    From the VIT internship: the system architecture and deployment decisions, including latency optimisation and where inference runs.

  2. Enhancing Drug Repositioning Through Collaborative Metric Learning

    IEEE, July 2024

    Collaborative metric learning for drug–disease association prediction, with better ranking performance on CTD benchmarks.

  3. AI Applications in Nutrition & Education

    IEEE Conference, 2024

    A survey of applied AI in nutrition and education that looks at deployment context and practical evaluation, not just benchmark accuracy.

About

I care less about how a model scores and more about whether it holds up.

I’m finishing a master’s in computer science at Rice, where I was also a teaching assistant for automata theory. Before that I studied electronics and computer engineering at VIT Chennai.

Most of my work sits between training a model and trusting it: fine-tuning it for a narrow task, finding where it fails, and keeping it reliable once it’s live. A benchmark score is where I start, not where I stop.

Outside of work, I support Chelsea, through the good seasons and the rest.

Toolkit

Languages & data

  • Python
  • Java
  • SQL
  • NumPy
  • Pandas
  • SciPy
  • SQLite
  • MongoDB

Machine learning

  • PyTorch
  • Hugging Face Transformers
  • TensorFlow
  • LoRA & quantization
  • LLMs & VLMs
  • Computer vision

LLM systems

  • Embeddings
  • RAG
  • Semantic search
  • Model evaluation

MLOps & cloud

  • AWS (EKS, EC2, S3)
  • GCP Cloud Run
  • Docker
  • FastAPI
  • MLflow
  • DVC
  • GitHub Actions
  • Prometheus & Grafana
  • Git

Contact

I graduate in December 2026 and I’m looking for full-time ML engineering roles, especially in LLM systems, model evaluation and MLOps. Email is the quickest way to reach me.

kaavinb7@gmail.com