Shashank Pandey

Lead AI Engineer at T-Mobile

Let’s connect

I build production AI systems, from LLM fine-tuning and retrieval to cloud infrastructure and reliable APIs.

Explore selected work →
Shashank Pandey
Scroll

3+ yrs production AI/ML · LLM platforms & RAG · multi-cloud MLOps · IEEE

// about

From applied research to production systems

I’m currently a Lead AI Engineer at T-Mobile, working remotely with an office in Seattle, WA. My experience spans LLM platforms, biomedical retrieval, forecasting, and production NLP.

I focus on evaluation, cost and latency, and maintainable software: reproducible experiments, clear APIs, and observable deployments.

// selected work

AI systems I’ve built

A closer look at platform engineering and applied research.

Multi-cloud LLM platform · Synergetics AI

LangTrain

Problem

Organizations need repeatable fine-tuning workflows across clouds without runaway GPU spend or ad-hoc scripts.

My contribution

Architected the platform, REST APIs, and microservices for multi-cloud LLM fine-tuning.

Architecture

Microservices and REST APIs; orchestrated training and reporting across AWS, Azure, and GCP.

Reported results

Reported ~40% reduction in GPU training cost; production use at Synergetics AI.

System overview · simplified
  1. 01REST APIs
  2. 02Training orchestration
  3. 03AWS · Azure · GCP

Python · FastAPI · AWS/Azure/GCP · GPU orchestration

Research · Jacob's Medicine · Jan 2025 – Dec 2025

GenAI Biomedical Pipeline

Problem

Research groups needed grounded retrieval and agentic workflows over diverse biomedical corpora.

My contribution

Designed biomedical GenAI systems and deployed retrieval and ML pipelines across 15+ research projects.

Architecture

LangChain and LangGraph orchestration, RAG with GPT-4 and Claude, FAISS and transformer stacks on Azure ML.

Reported results

~35% retrieval efficiency gains; ~28% improvement in semantic search accuracy; ~45% inference latency reduction via distillation/pruning with quality held.

System overview · simplified
  1. 01Biomedical corpora
  2. 02FAISS retrieval
  3. 03LangGraph · GPT-4 / Claude

LangChain · LangGraph · RAG · Azure ML

More projects

// experience

CURRENT

Lead AI Engineer

T-Mobile

Enterprise AI agents and telecommunications data integration on Azure, with a focus on inference efficiency and application observability.

May 2026 - Present
Remote — Seattle, WA office

  • Optimized LLM prompts and agent workflows, reducing token consumption by 25% and inference costs by 20% while improving GPU and provisioned throughput unit (PTU) utilization.
  • Managed Azure Cosmos DB and Azure ML infrastructure, reducing query latency by 40%; used Azure Application Insights to monitor application performance and failures.
  • Developed 5 Model Context Protocol (MCP) servers and integrated 10 APIs to connect AI agents with enterprise tools and telecommunications data sources.
PRODUCTION

AI/ML Engineer – Software Developer

Synergetics AI

Production LLM platform, multi-cloud GPU orchestration, APIs and full-stack delivery.

Dec 2025 - May 2026
Remote, USA

  • Architected LangTrain, a full-stack multi-cloud LLM fine-tuning platform with RESTful APIs and microservices across AWS, Azure, and GCP, reducing GPU training costs by 40%.
  • Engineered scalable data and reporting applications using Python, FastAPI, and Superset with sub-second latency for executive and operational reporting.
  • Implemented an AI agent marketplace application (4,800+ LOC) with unit and integration testing, code reviews, Git workflows, and CI/CD, streamlining agent onboarding.
More about this role
  • Defined RESTful API design and service integration patterns; delivered documented, maintainable code and drove adoption across cross-functional teams.
  • Deployed full-stack features with React/Next.js frontends, Node.js and Python backends, and PostgreSQL/cloud databases; designed schemas and SQL for reporting and agent state.
  • Optimized inference footprint by 35% through prompt engineering, fine-tuning frameworks, and model compression and quantization while preserving output quality.
RESEARCH

ML Engineer – GenAI

Jacob's Medicine and Biomedical Sciences

GenAI systems for research: RAG, LangGraph, Azure ML, compression and latency work.

Jan 2025 - Dec 2025
Buffalo, USA

  • Designed GenAI biomedical systems using LangChain, LangGraph, GPT-4, Claude, and RAG across 15+ research projects, improving data retrieval efficiency by 35%.
  • Deployed NLP systems with BERT, T5, and FAISS; developed feature engineering and data augmentation pipelines that improved semantic search accuracy by 28%.
  • Automated end-to-end ML pipelines on Azure ML with PyTorch for preprocessing, training, hyperparameter tuning, and evaluation, reducing training cycles by 40%.
More about this role
  • Reduced inference latency by 45% through knowledge distillation and weight pruning on production language models while preserving ~95% accuracy.
PRODUCTION

ML Engineer – Data Science

Baldwin Richardson Foods

Forecasting and anomaly detection at scale; MLOps on AWS and Databricks.

Aug 2024 - Dec 2024
New York, USA

  • Delivered anomaly detection and time-series forecasting models using XGBoost, LSTM, and Prophet on 14+ years of sales data, improving prediction accuracy by 30% and enabling $1.5M in annual cost savings.
  • Automated feature engineering and model evaluation pipelines on AWS SageMaker and Databricks with experiment tracking for faster model iteration.
  • Deployed model serving infrastructure with Docker, Kubernetes, and GitHub Actions CI/CD, reducing deployment effort by 70% and enabling production monitoring and refresh.
More about this role
  • Developed efficient SQL in PostgreSQL and Oracle for high-volume operational analytics; built Tableau and Streamlit dashboards for cross-functional stakeholders.
PRODUCTION

Machine Learning Engineer

KPIT Technologies

Production NLP over high-volume tickets; IoT analytics; TensorFlow Serving and MLOps.

Jun 2021 - Jul 2023
India (Remote)

  • Deployed production NLP pipelines for text classification, entity extraction, and incident categorization across 50K+ tickets using BERT and PyTorch, achieving 0.88 precision and 0.85 recall and reducing manual review time by 60%.
  • Implemented anomaly detection and feature-engineering frameworks on large-scale automotive IoT datasets using PySpark and AWS SageMaker, enabling earlier detection of equipment failures.
  • Built MLOps workflows with TensorFlow Serving for model monitoring, versioning, and refresh management in production environments.
More about this role
  • Standardized ML pipeline architecture with cross-functional teams; introduced validation, A/B testing, code reviews, and deployment best practices.
MISSION

Data Scientist – Software Intern

ISRO - Indian Space Research Organisation

Spectral DL and peak detection for space-like conditions; distributed Spark; mission reliability.

Sep 2021 - Aug 2022
Bangalore, India

  • Developed deep learning classifiers and automated peak-finding algorithms on satellite telemetry and laser-induced plasma spectra using PyTorch, improving anomaly detection accuracy by 32% (IEEE Conference 2022).
  • Deployed real-time model serving and early-warning analytics for signal degradation detection, achieving 99.9% system reliability on mission-critical infrastructure.
  • Scaled distributed Spark pipelines for high-volume data storage and processing, delivering 10× throughput with automated monitoring and model refresh workflows.
More about this role
  • Rolled out CI/CD and monitoring workflows across 15+ cross-functional teams to standardize data pipelines for mission-critical reliability.

// volunteering

[VOLUNTEER]

Data Engineer - Generative AI

Community Dreams Foundation

Science and Technology

Feb 2025 - Oct 2025
9 mos

  • Built Azure data pipelines and ETL workflows with Data Factory, Databricks, Blob Storage, and ADLS for analytics and AI/ML workloads.
  • Integrated Generative AI workflows and automated data validation to support nonprofit service delivery.
  • Documented data models and workflows to support contributor onboarding and ongoing maintenance.

// freelance

FREELANCE

Machine Learning & Data Science Expert (Freelance)

AfterQuery Experts

Mar 2026 - Present
Remote

  • Build, train, and evaluate machine learning models for client business applications including forecasting, classification, and recommendation use cases.
  • Develop predictive models using supervised and unsupervised learning techniques; deliver interpretable outputs and reports for stakeholders.
  • Optimize model performance through hyperparameter tuning and feature engineering; document methodologies, assumptions, and technical approaches for reproducibility and handoff.
  • Collaborate with clients to define problem scope, success metrics, and delivery timelines for ML and analytics engagements.

// skills

Telecom AI & Integration

Telecommunications data integration · Model Context Protocol (MCP) · Enterprise APIs · AI agent workflows · Azure Cosmos DB · Azure Application Insights · GPU / PTU optimization

Programming Languages

Python · SQL · R · C++ · Java · JavaScript

Frontend

React.js · Next.js · HTML · CSS · Tailwind CSS

Backend

Node.js · Express · Nest.js

ML & AI

PyTorch · TensorFlow · Scikit-learn · Pandas · PySpark · NLP (Transformers, BERT) · LLMs (GPT-4, LLaMA, Gemini) · LLM Fine-tuning · Prompt Engineering · CNNs · LSTMs · XGBoost · Anomaly Detection · Forecasting

ML Infrastructure

Feature Engineering · Model Training · Model Serving · Model Monitoring · Hyperparameter Tuning · Model Compression

Cloud & MLOps

AWS SageMaker · Azure ML · GCP · Databricks · Docker · Kubernetes · CI/CD · GPU Optimization

Visualization

Tableau · Streamlit · Power BI

// certifications

AWS Machine Learning SpecialtyMicrosoft Power BIIBM AI Engineering Professional CertificateAWS Certified Cloud Practitioner

// education

University at Buffalo

Masters, Data Science

Jul 2023 - Dec 2024 · USA · GPA: 3.8

International Institute of Information Technology-Bangalore

Associate's degree, Data Science

Dec 2021 - Jan 2023 · India

M. S. Ramaiah University of Applied Sciences

Bachelor of Technology, Electronics and Communication Engineering

Aug 2018 - May 2022 · India

// publications

An Automatic Peak Finding and Fitting Aspects of Laser Induced Plasma Spectra acquired in High Vacuum: Tradeoff Simulations and Statistics

IEEE

Problem: overlapping peaks in laser-induced plasma spectra hurt elemental ID in vacuum conditions relevant to space science. Method: simulation-driven trade-offs and parameter tuning for peak finding. Outcome: clearer guidance for mitigating overlap and improving detection fidelity.

Full abstract

Laser Induced Plasma Spectroscopy (LIPS) is a promising spectrochemical analytical method for rapid analysis of multi-element samples, and has become a potential field of both fundamental and exploratory research including the space science in recent times. Although the LIPS technique is highly versatile, its element detection capability at times is intriguing due to spectral peak overlapping that can hamper the elemental detection accuracy. This paper presents details on executed trade-off simulations and optimization of algorithm parameters that may aid for effective mitigation of spectral overlapping issues of LIPS spectra and precise peak finding. Four pelletized samples were used to acquire spectra in high vacuum (≤ 5×10⁻⁶ mbar) environment to mimic space-like conditions.

View on IEEE Xplore →

// writing

In the ever-evolving landscapes of retail and e-commerce, managing demand volatility, intricate supply chains, and unpredictable customer behavior has become a data-intensive challenge. For global titans like Walmart and Amazon, solving this challenge at scale means embracing advanced time series forecasting techniques rooted in both academic research and operational excellence.

Medium

TensorRT is NVIDIA’s SDK for optimizing and deploying deep learning models for inference. It takes trained models from frameworks like PyTorch, TensorFlow, or ONNX and tunes them to run faster, leaner, and smarter on NVIDIA GPUs.

Medium

Imagine asking your smart assistant about yesterday’s football match, and it confidently describes a game from 2022. Frustrating, right? This is where RAG (Retrieval-Augmented Generation) comes in — it’s like giving your AI a real-time sports ticker instead of relying on old memories.

Medium