AI/ML engineer building production ML systems and LLM tooling — from autonomous-driving perception to LLM-powered analytics pipelines.
AI Engineer with 3+ years building production LLM systems, multi-agent pipelines, and inference optimization engines. Currently completing an MS in Applied Data Science at USC while working as an AI Engineer at USC Marshall School of Business and USC AutoDrive Lab.
Recent work includes building Inferno — a KV cache quantization and continuous batching engine benchmarked on Tesla T4 GPU — a multi-agent toxicity evaluation pipeline at Convexia (YC'25), and an agentic VLM + OCR document extraction system at USC Marshall. Won 1st place among 100+ teams at the Miro × Kiro LA Hackathon building a multi-agent AI restaurant recommendation system.
Open to full-time AI Engineer and ML Engineer roles starting December 2026.
Master of Science in Applied Data Science
Focused on advanced machine learning, statistical modeling, and scalable AI systems for real-world applications.
Bachelor of Technology in Artificial Intelligence and Data Science
Comprehensive undergraduate program covering AI fundamentals, machine learning algorithms, data structures, and blockchain technology. Specialized coursework in deep learning, computer vision, and natural language processing.
From-scratch LLM inference optimization engine implementing INT8 KV cache quantization (per-tensor and per-channel) and a continuous batching scheduler on top of HuggingFace Transformers. Benchmarked on Tesla T4 GPU — 2x KV cache compression, 46/46 tests passing. All benchmark claims sourced to reproducible JSON outputs.
Extract structured company records from scanned historical directory pages. CLI processes a single page image or a full PDF with two-stage triage (lightweight scan, then full extraction), multi-variant Tesseract OCR, staff-region OCR, and optional LLMWhisperer VLM fallback when required fields are empty. Exports spreadsheet-friendly CSV, XLSX, and JSON.
Agentic AI-powered educational platform with RAG-based Q&A, quiz generation, and real-time chat using Cohere embeddings, Pinecone, LangChain, GPT-4, and FastAPI, supporting semantic retrieval across 10,000+ pages of textbook content.
Scalable multi-level text retrieval framework integrating Gaussian Mixture Models, GPT-4-turbo, BM25, and semantic search, enabling granular content extraction from 300+ page academic textbooks stored in ChromaDB, improving re-ranking precision by 30%.
End-to-end histopathology image classification pipeline using MobileNetV2 at 92% accuracy, with Grad-CAM interpretability and DCGAN data augmentation.
End-to-end toxicity evaluation pipeline at Convexia (YC'25) integrating 6 ML models with MLflow tracking, SHAP feature visualization, and confidence/disagreement detection for 100% reproducibility.
Hybrid coach selection engine at Easley Dunn Productions combining rule-based constraints with AI scoring across 15+ attributes, improving team matching accuracy by 25% and cutting lineup imbalance by 40%.
Perception, motion-prediction, and planning models at USC AutoDrive Lab; deployed transformer-based generative AI on CUDA clusters (4× faster training) and Deep RL (PPO/SAC) reaching 0.4 m mean positional deviation over 500+ closed-loop runs.
This publication explores the implementation impact of AI across developing industries, highlighting practical adoption patterns, opportunities, and challenges in real-world systems.
This work examines the Blockchain-Based Election Conducting and Management System (BCECMS) as a framework to improve election integrity, transparency, and process trust through auditable digital workflows.
The best way to reach me is directly:
Connect on LinkedIn or explore my work on GitHub below.