Open to new opportunities

Building enterprise-grade AI systems
with RAG, Agentic AI, evaluation
frameworks, observability, and cloud deployment.

AI Engineer with 13+ years of total experience — 3 years in AI and 10 years in enterprise software engineering. Building enterprise-grade AI systems with RAG, Agentic AI, evaluation frameworks, observability, and cloud deployment.

10+
Years in Software Engineering
3+
Years in AI
2
Live Deployed Projects
// live system
Enterprise
Advanced RAG
Guardrails9 layers
Cache tiers5 active
Test coverage184 tests
Eval score0.97 faithfulness
InfraAWS ECS Fargate
▶ View Live Demo
// portfolio

AI Projects

Two flagship projects across two clouds and two architectural patterns — Enterprise RAG on AWS and multi-agent clinical intelligence on GCP.

● LIVE
🤖 Enterprise RAG · IT Operations

Enterprise Advanced RAG — Kubernetes IT-Operations Copilot

Production-deployed enterprise RAG system for Kubernetes IT operations. Hybrid retrieval with cross-encoder reranking, CRAG, HyDE, Self-RAG reflection, Text-to-SQL with human-in-the-loop approval, 9-layer guardrails, sqlglot safety validation, semantic caching, and admin observability — all in one production workflow. Auto/Manual experience modes with preset query library.

Hybrid + RerankCRAG HyDESelf-RAG Text-to-SQL + HITLsqlglot Safety 9-Layer GuardrailsQdrant FastAPIAWS ECS Fargate
◌ IN BUILD
🧬 Multi-Agent · Clinical Trial · Research Risk

Multi-Agent Clinical-Trial Intelligence & Research-Risk Assistant

Supervisor-orchestrated system of 6 specialist agents that autonomously audit 400K+ ClinicalTrials.gov studies for broken promises, missing results, sponsor credibility, cross-study patterns, safety gaps, and timeline deviations. Three-tier LangMem memory (Episodic · Procedural · Semantic) persists knowledge across sessions — human rejections permanently update procedural memory so agents learn from every review cycle. Per-agent HITL confidence thresholds (0.55–0.70) gate low-confidence signals to a human review queue with a learning loop that injects new rules back into agent context. Deployed serverless on GCP Cloud Run.

LangGraph Supervisor Agent LangMem · 3-tier Memory 6 Specialist Agents Human-in-the-Loop FastAPI GCP Cloud Run Cloud SQL · PostgreSQL ClinicalTrials.gov API PubMed eUtils API OpenAI GPT-4o LangSmith
Data scale
400K+
studies audited
Agents
6 specialist
parallel execution
Cloud
GCP
serverless · $35–65/mo
◌ IN BUILD
🧠 Deep Learning · CV

CNN Image Classifier

50-class image classifier built from scratch in PyTorch — no pretrained backbone. Custom architecture, data augmentation pipeline, checkpoint management, and training loop with learning-rate scheduling. Benchmarked against transfer-learning baselines.

PyTorchCNN Data AugmentationAWS ECS
◌ IN BUILD
🎨 Deep Learning · Generative

GAN Face Generation

Generative adversarial network for face synthesis — built from scratch in PyTorch. Custom discriminator/generator architecture, training stability techniques (label smoothing, gradient penalty), and mode collapse mitigation. Evaluated on FID score and visual coherence.

PyTorchGAN Generative AIComputer Vision
○ PLANNED
📊 Evaluation · RAG

LLM Evaluation Framework

Production-grade evaluation harness for RAG systems and LLM pipelines. RAGAS metrics, deterministic route/source gates, per-question thresholds, sealed holdout sets, and impossible-score guards. Designed to prevent evaluator confidence from hiding product quality issues.

RAGASDeepEval PythonLangGraph
○ PLANNED
🤖 Agentic AI · Tool Use

Tool-Using Agent System

LangGraph agent that orchestrates specialized ML models and APIs as tools — decides which tool to invoke, interprets structured results, and synthesizes evidence into grounded responses. Demonstrates the "agent orchestrates specialized models" production pattern.

LangGraphTool Use PyTorchFastAPI

// technical writing

Blog & Write-ups

Deep dives into the architecture decisions, evaluation challenges, and engineering trade-offs behind these production systems.

🤖 RAG · Architecture
Building an Enterprise-Grade Advanced RAG System: What It Took Beyond the Demo
How I designed a production-grade RAG system — covering hybrid retrieval, intent routing, CRAG, HyDE, Self-RAG, 9-layer guardrails, safe Text-to-SQL with human approval, and an evaluation framework that cannot hide failure.
2026 · 15 min read ↗ Read on Medium
COMING SOON 📊 Evaluation · LLM
RAGAS Evaluation: Making LLM Quality Actually Measurable
Why high aggregate scores can hide serious quality failures — and how I built an evaluation harness with deterministic gates, sealed holdout sets, impossible-score guards, and per-question thresholds that cannot be gamed.
2026 · Upcoming ~10 min read
COMING SOON 🔒 SQL · Safety
Text-to-SQL with Human-in-the-Loop Approval
Generated SQL is untrusted code. A walkthrough of the full safety pipeline — schema-scoped generation, sqlglot AST validation, read-only bounds, stale-approval rejection, and human approval — and why each layer matters.
2026 · Upcoming ~8 min read

// background

Experience

2024 — Present
AI Engineer
3 Years · AI Engineering · aiscientificlabs.com
Building and deploying production AI systems — Enterprise RAG pipelines, multi-agent orchestration with LangGraph, hybrid retrieval, Text-to-SQL with human-in-the-loop approval, 9-layer guardrails, and LLM evaluation frameworks. Deployed on AWS ECS Fargate with full CI/CD and CloudWatch observability. LangGraph, Qdrant, FastAPI, Streamlit.
2014 — 2024
Senior Software Engineer
10 Years · SAP Analyst · Salesforce Admin · Project Manager · Automation Testing · Performance Testing
A decade of enterprise software engineering across diverse domains — SAP systems analysis, Salesforce administration and customisation, project management, automated testing frameworks, and performance testing at scale. Delivered complex cross-functional programmes across global enterprise environments.

// tech stack

Production-Grade Tools

Every layer used in these deployed systems — chosen for production, not notebooks.

AI / ML
PyTorchLangGraph LangChainTransformers (HF) OpenAI APIDeepEval RAGAS
Vector / Search
QdrantBM25 Hybrid Cross-Encoder Reranker HyDECRAG
Backend / API
PythonFastAPI StreamlitPydantic Docker
AWS Cloud
ECS FargateALB ACM / HTTPSECR CloudWatchRoute 53IAM
Evaluation
RAGASDeepEval W&B / MLflow Gold-Set QADistractor Testing

Let's build production AI together

Got an interesting AI architecture challenge? Let's talk.