Aryan Maheshwari
Building LLM inference systems and production AI pipelines

Hey there — I’m Aryan Maheshwari

AI/ML engineer building production ML systems and LLM tooling — from autonomous-driving perception to LLM-powered analytics pipelines.

About Me

AI Engineer with 3+ years building production LLM systems, multi-agent pipelines, and inference optimization engines. Currently completing an MS in Applied Data Science at USC while working as an AI Engineer at USC Marshall School of Business and USC AutoDrive Lab.

Recent work includes building Inferno — a KV cache quantization and continuous batching engine benchmarked on Tesla T4 GPU — a multi-agent toxicity evaluation pipeline at Convexia (YC'25), and an agentic VLM + OCR document extraction system at USC Marshall. Won 1st place among 100+ teams at the Miro × Kiro LA Hackathon building a multi-agent AI restaurant recommendation system.

Open to full-time AI Engineer and ML Engineer roles starting December 2026.

Education

🏛️ University of Southern California

Jan 2025 – Dec 2026 Los Angeles, California

Master of Science in Applied Data Science

Focused on advanced machine learning, statistical modeling, and scalable AI systems for real-world applications.

🏫 K.J. Somaiya Institute of Technology

2020 – 2024 Mumbai, India

Bachelor of Technology in Artificial Intelligence and Data Science

Comprehensive undergraduate program covering AI fundamentals, machine learning algorithms, data structures, and blockchain technology. Specialized coursework in deep learning, computer vision, and natural language processing.

Experience

🔬 AI Engineer — USC Marshall School of Business

Dec 2025 – Present Los Angeles, CA
LLM ETLAgentic NLPVLMOCRDocument AI

🤖 AI Engineer — Convexia (YC'25)

Mar 2025 – Sep 2025 San Francisco, CA
MLflowSHAPMLOpsCI/CDPython

🎮 Machine Learning Engineer — Easley Dunn Productions, Inc

Jun 2025 – Sep 2025 Los Angeles, CA
Recommendation SystemsHybrid MLTeam LeadershipPython

🚗 Machine Learning Engineer — USC Autodrive Lab

Jun 2025 – Present Los Angeles, CA
Deep RLPPOSACCUDAAutonomous Driving

🤖 AI Engineer — AGIE AI

Jan 2024 – Nov 2024 Mumbai, India
DialogflowVertex AIGCPEmbeddingsRetrieval

🧪 AI Engineer — Dawn Digitech

Jan 2023 – Dec 2023 Mumbai, India
LLMsSentiment AnalysisBERTVADERNLP

Achievements

🥇 1st Place — MiroxKiro LA Hackathon

1st among 100+ teams

Architected a multi-agent AI system with LLM-based reasoning, retrieval pipelines, and geospatial intelligence to deliver personalized recommendations and spatial experiences.

Multi-Agent SystemsLLM ReasoningRAGGeospatial AI

Projects

Inferno — LLM Inference Optimization Engine

From-scratch LLM inference optimization engine implementing INT8 KV cache quantization (per-tensor and per-channel) and a continuous batching scheduler on top of HuggingFace Transformers. Benchmarked on Tesla T4 GPU — 2x KV cache compression, 46/46 tests passing. All benchmark claims sourced to reproducible JSON outputs.

KV CacheINT8 QuantizationContinuous BatchingPyTorchCUDAHuggingFaceTesla T4

Zage OCR — Document Extraction Pipeline

Extract structured company records from scanned historical directory pages. CLI processes a single page image or a full PDF with two-stage triage (lightweight scan, then full extraction), multi-variant Tesseract OCR, staff-region OCR, and optional LLMWhisperer VLM fallback when required fields are empty. Exports spreadsheet-friendly CSV, XLSX, and JSON.

OCRTesseractPDF PipelineVLMLLMWhispererPythonDocument AI

🎓 EduMate.ai

Agentic AI-powered educational platform with RAG-based Q&A, quiz generation, and real-time chat using Cohere embeddings, Pinecone, LangChain, GPT-4, and FastAPI, supporting semantic retrieval across 10,000+ pages of textbook content.

RAGCoherePineconeLangChainGPT-4FastAPI

📚 HieQue

Scalable multi-level text retrieval framework integrating Gaussian Mixture Models, GPT-4-turbo, BM25, and semantic search, enabling granular content extraction from 300+ page academic textbooks stored in ChromaDB, improving re-ranking precision by 30%.

GMMGPT-4-turboBM25ChromaDBSemantic Search

🔬 HistoHelp

End-to-end histopathology image classification pipeline using MobileNetV2 at 92% accuracy, with Grad-CAM interpretability and DCGAN data augmentation.

MobileNetV2Grad-CAMDCGANStreamlitTensorFlow

🧪 Toxicity Evaluation Pipeline

End-to-end toxicity evaluation pipeline at Convexia (YC'25) integrating 6 ML models with MLflow tracking, SHAP feature visualization, and confidence/disagreement detection for 100% reproducibility.

MLflowSHAPMLOpsCI/CDPython

🏅 Coach Selection Engine

Hybrid coach selection engine at Easley Dunn Productions combining rule-based constraints with AI scoring across 15+ attributes, improving team matching accuracy by 25% and cutting lineup imbalance by 40%.

Machine LearningPythonSports AnalyticsScikit-learn

🚗 Autonomous Driving Perception & Planning

Perception, motion-prediction, and planning models at USC AutoDrive Lab; deployed transformer-based generative AI on CUDA clusters (4× faster training) and Deep RL (PPO/SAC) reaching 0.4 m mean positional deviation over 500+ closed-loop runs.

PyTorchDeep RL (PPO/SAC)CUDATransformersAutonomous Driving

Technical Skills

Languages

PythonJavaC++TypeScriptJAXSQLNoSQLReact

Machine Learning

Multi-Agent SystemsLLM Fine-TuningDeep LearningRAGVector DatabasesReinforcement Learning

ML Systems & Infrastructure

LLM EvaluationPrompt EngineeringMCPOrchestrationCUDAApache SparkMLflow

Cloud & DevOps

AWS (S3, OpenSearch)GCP (Vertex AI, Dialogflow)DockerCI/CD

Research

🤖 Role of AI in Developing Industries

Published 2024 JETIR

This publication explores the implementation impact of AI across developing industries, highlighting practical adoption patterns, opportunities, and challenges in real-world systems.

🔗 Exploring the Implementation of BCECMS for Election Security and Transparency

Published 2024 IEEE Xplore

This work examines the Blockchain-Based Election Conducting and Management System (BCECMS) as a framework to improve election integrity, transparency, and process trust through auditable digital workflows.

Contact

The best way to reach me is directly:

📧 aamahesh@usc.edu

Connect on LinkedIn or explore my work on GitHub below.