Open to ML Engineer, AI Engineer & Data Scientist roles

Hi, I'm

Kiran Babu Athina

Machine Learning & AI Engineer

I build

From agentic RAG on GPU clusters to PySpark pipelines moving 100M+ records a day, I ship AI that holds up in production, and every system comes with an eval set and a baseline.

Portrait of Kiran Babu Athina

Rigorous evaluation, production-grade AI

I'm a Machine Learning and AI Engineer with four years of experience building and shipping production ML across telecom, finance, healthcare, and agricultural research.

At Texas A&M's AISLS Lab I build agentic RAG systems over knowledge graphs and fine-tune open LLMs like Gemma and LLaMA on local GPU infrastructure. Before that, at Tata Consultancy Services, I engineered PySpark pipelines, NLP classifiers, and multimodal QC models running on Docker and Kubernetes.

I hold an M.S. in Data Science from Texas A&M (3.9 GPA). I don't call a model done until it has a held-out eval set, a baseline to beat, and a clear recommendation stakeholders can act on.

4+
Years building production ML
98.7%
Recall@5, agentic KG-RAG
100M+
Records/day through PySpark ETL
3.9
GPA, M.S. Data Science

Where I've worked

Research Assistant

Jun 2025 – Present

Texas A&M AISLS Lab · College Station, TX

  • Achieved 98.7% Recall@5 across 3 domain-specific knowledge graphs with an agentic RAG system on TAMU HPC (2× NVIDIA RTX 6000), using a query-routing layer over 3 fine-tuned KG retrievers.
  • Cut general-purpose cloud API calls to zero for precision nutrition queries by fine-tuning Gemma 4 12B with LoRA on food-systems corpora aligned with SDG 2.
  • Reached 95% Hit@5 and 0.925 MRR on a 1,000-question expert-curated eval set with a FastAPI RAG backend on Azure App Service (hybrid BM25 + FAISS over 120,000+ chunks).
  • Built a chunk-level scoring and citation pipeline inside the inference server: 86% source precision, 95% source recall.
  • Eliminated cold-start 503s, with zero failed session initializations, through a React chat frontend with race-safe bootstrapping, streaming UX, and path-triggered CI/CD on Azure.
Agentic RAGKnowledge GraphsLoRA FastAPIFAISSReactAzure

Project Lead

Aug 2025 – May 2026

Aggie Research Program · Texas A&M University

  • Led teams of 3 undergraduate researchers per project, owning end-to-end execution aligned with faculty research goals.
  • Scoped projects, broke research objectives into weekly tasks, and mentored students in applied ML, data analysis, and research practice.
  • Served as primary point of accountability, reporting progress, risks, and results to the supervising professor weekly.
LeadershipMentoringResearch Management

Graduate Research Assistant

Jan 2024 – May 2025

Texas A&M Department of Animal Science · College Station, TX

  • Eliminated Azure OpenAI calls for livestock terminology queries by fine-tuning LLaMA 3 with LoRA/PEFT on a proprietary animal science corpus.
  • Delivered grounded, citation-backed answers with a RAG pipeline on Azure AI Search + Azure OpenAI, validated by faculty review across 3 semesters.
  • Built a sensor-driven digital twin (MQTT + Mesa agent-based model) for minute-level per-animal energy tracking and monthly methane forecasting.
LLaMA 3PEFTAzure OpenAI Mesa ABMMQTT

Machine Learning Engineer

Aug 2021 – Jul 2023

Tata Consultancy Services · Hyderabad, India

  • Cut loan processing time by 30% with classification models on Azure ML using AutoML selection and statistical feature engineering.
  • Eliminated 30% of manual hardware log triage and sped up issue resolution by 18% with a TF-IDF + Random Forest NLP pipeline over 500,000+ daily telecom logs.
  • Reached 92% pass/fail accuracy on hardware label QC across 10,000+ devices/month with a multimodal ResNet-18 + BERT verification system.
  • Processed 100M+ sensor records/day for fault-prediction models with a PySpark ETL framework on Hadoop.
  • Reduced post-deployment performance decay by 22% with drift dashboards and retraining triggers, and cut deployment turnaround by 30% with Docker + Kubernetes auto-scaling.
PySparkHadoopAzure ML BERTResNetDockerKubernetes

Machine Learning Engineer Intern

May 2021 – Jul 2021

Tata Consultancy Services · Hyderabad, India

  • Achieved 98.5% fraud classification accuracy on KYC anomaly detection with a multimodal PyTorch model fusing 1D CNN audio and 2D CNN document-image embeddings.
  • Reduced onboarding and handoff time by 25% by writing end-to-end documentation of model architectures, assumptions, and evaluation results.
PyTorchCNNsMultimodalFraud Detection

My toolkit

Languages

PythonSQLJavaScript RC

ML / AI

LLMsRAGAgentic AI Knowledge GraphsLoRAQLoRA PEFTTransformersVision-Language Models CLIPPrompt EngineeringPyTorch TensorFlowHugging FaceScikit-learn LangChainLangGraphFAISS NLTKXGBoostLightGBM ARIMA / SARIMA / LSTM

Cloud & HPC

Azure AI SearchAzure AI FoundryAzure OpenAI Azure MLAzure App ServiceAzure Static Web Apps Azure Blob StorageAWS S3AWS EC2 AWS SageMakerNVIDIA GPU Clusters

Tools & Data

DockerKubernetesCI/CD Git / GitHubFastAPIReact PySparkHadoopStreamlit TableauPostgreSQLMySQL MongoDBRedisMQTT

Things I've built

Research

Agentic RAG over Domain Knowledge Graphs

Agentic retrieval system on TAMU HPC with a query-routing layer that sends each question to one of 3 fine-tuned knowledge-graph retrievers.

98.7% Recall@5 on held-out eval
Agentic AIKnowledge GraphsQuery RoutingTAMU HPCRTX 6000
Research

Source-Grounded RAG Platform on Azure

FastAPI backend plus React chat frontend indexing 120,000+ document chunks with hybrid BM25 + FAISS retrieval and top-5 source traceability on every answer.

95% Hit@5 · 0.925 MRR
FastAPIBM25FAISSReactAzure
Healthcare AI

Retinal Disease Diagnosis with Vision-Language Models

Fine-tuned MedGemma 4B on paired fundus images and clinical notes for 12-class retinal disease classification, validated on 1,200+ expert-annotated images.

91% diagnostic accuracy
MedGemma 4BVLMsFine-tuningMedical Imaging
Deep Learning

CLIP-style Multimodal Contrastive Learning

Dual-encoder ResNet-50 + DistilBERT model trained with NT-Xent and iSogCLR losses and AdamW/RAdam optimizers, benchmarked against baseline CLIP.

+17% retrieval (MSCOCO) · +14% zero-shot (ImageNet)
PyTorchResNet-50DistilBERTContrastive Learning
NASA Space Apps 2025

Exoplanet Discovery

Ensemble of Random Forest, XGBoost, LightGBM, Gradient Boosting, and MLP classifying confirmed exoplanets, candidates, and false positives from Kepler and TESS data.

96.64% classification accuracy
XGBoostLightGBMScikit-learnStreamlit
Research

Livestock Digital Twin

Sensor-driven agent-based simulation that ingests live MQTT sensor streams to model per-animal energy and methane dynamics.

Minute-level energy tracking · monthly CH₄ forecasts
PythonMesa ABMMQTTSimulation

Research output

Indexed on Google Scholar · Texas A&M University · Artificial Intelligence Scholar profile
2025
An agent-based framework of cattle value discovery system for precision nutrient requirements and utilization prediction of beef cattle

K. Kaniyamattam, V. Kulangara-Veettil, P. Kundu, K. B. Athina, L. O. Tedeschi

CABI Digital Library · Vol. 16, Issue 3, p. 449 · September 2025

Translates the deterministic Cattle Value Discovery System (CVDS) into an agent-based model, simulating individual cattle as autonomous agents with their own body weight, body condition score, and metabolic efficiency to predict daily gain, days to finish, and carcass composition.

0 Citations

Academic background

Texas A&M University

M.S. in Data Science

Aug 2023 – May 2025 · College Station, TX

GPA 3.9 / 4.0

Gayatri Vidya Parishad College of Engineering

B.Tech. in Electronics & Communication Engineering

Aug 2017 – Jul 2021

First Class with Distinction

Certifications

NVIDIA DLI Certified Generative AI Fundamentals PCAP: Programming Essentials in Python CCNA: Introduction to Networks SailPoint Identity Security Leader Intermediate SQL Queries

Let's build something

I'm open to ML Engineer, AI Engineer, and Data Scientist roles. Based in Austin, TX. The fastest way to reach me is email.