AgenticBI — AI-Powered BI Copilot Platform
Jan 2026 – Present
- • Architected a BI-tool-agnostic event-driven AI copilot enabling multimodal (voice and text) interaction with business data, automating end-to-end BI workflows — query generation, execution, dashboarding, and insight explanation.
- • Using a multi-agent LangGraph microservice system with self-correcting LLM feedback loops, MCP server for API-based tool orchestration, GraphRAG retrieval for schema-aware context grounding, and a Whisper-powered voice-input pipeline. Designed layered guardrails for prompt injection defence and secure SQL execution.
SLM Agentsllama-3.3Qwen-32BGroqMCPA2AChromaDBVizroLangChainLangGraphDockerGit
LLM Advisor – Trust-Based Model Recommender for Open-Source LLMs
Dec 2025 - Jan 2026
- • Built a two-stage parallel evaluation and scoring pipeline that profiles open-source LLMs: Llama‑3 8B, Qwen 2.5 7B, Kimi K2 on toxicity, bias, truthfulness, and safety to generate reproducible trustworthiness profiles for ethical AI deployment.
- • Developed an LLM-powered advisor using Mixtral‑8x7B that ingests model profiles and user priorities to recommend suitable LLMs, explaining trade-offs between safety, fairness, and factual reliability for real-world deployments.
Groq-hosted LLMsHuggingFace DatasetsDetoxifysentence-transformersStreamlitPlotlyPython
Hierarchical Reinforcement Learning for Portfolio Management
Aug 2025 – Dec 2025
- • Developed a three-tier Hierarchical RL system for multi-asset portfolio optimization, integrating market data and Reddit sentiment to achieve stronger risk-adjusted returns and reduced drawdowns across 3 years of back testing.
- • Trained and aggregated SAC, PPO, DDPG, and TD3 agents on FinBERT sentiment and Yahoo Finance pipelines, combining them through meta- and super-agent layers to outperform equal-weighted and single-agent baselines.
PyTorchgymnasiumstable-baselines3transformersFinBERTyFinanceReddit APIdata preprocessing
Hardware-Accelerated LLM Inference Engine
Oct 2025 – Dec 2025
- • Built an optimized inference runtime for Llama-3-8B utilizing TensorRT-LLM, applying W8A8 PTQ via NVIDIA AMMO to measure a 52% reduction in memory footprint, enabling high-throughput deployment on 40GB NVIDIA A100 GPU.
- • Replaced native PyTorch attention mechanisms by compiling the engine with FlashAttention-2 and paged KV cache enabled, optimizing memory fragmentation to increase the maximum concurrent batch size by 3x.
- • Profiled the C++ runtime execution utilizing NVIDIA Nsight Systems to identify and minimize CPU-GPU kernel dispatch overhead, achieving a throughput of 85 tokens/sec, a 2.1x performance increase compared to a baseline vLLM FP16.
C++CUDATensorRT-LLMNVIDIA AMMOFlashAttention-2NVIDIA NsightvLLMPost-Training Quantization
Graph-of-Thought Trajectory Prediction with LightEMMA
June 2025 – Aug 2025
- • Integrated Graph-of-Thought reasoning into the LightEMMA Vision- Language-Model trajectory prediction framework, adding multi-intent generation, branch scoring, and refinement to improve motion forecasting quality on nuScenes dataset compared to Chain-of-Thought baselines using Claude Sonnet 4.5 and Google Gemini Flash 2.0.
Vision-Language ModelsTransformersHugging FacenuScenes DevKitOpenCVSciPyScikit-learn
Real-Time Edge Perception Pipeline
Aug 2025 - Oct 2025
- • Accelerated a PointPillars 3D object detection model for a 15W NVIDIA Jetson Orin Nano by exporting to ONNX and compiling an FP16 TensorRT engine, achieving a 3.4x inference latency speedup over an ONNX Runtime baseline.
- • Engineered a native C++ multithreaded preprocessing pipeline utilizing a 4-worker thread pool to execute LiDAR point cloud voxelization concurrently, reducing CPU-bound data preparation latency by 55%.
- • Designed a zero-copy memory architecture using CUDA pinned memory to eliminate redundant CPU-GPU data transfers, sustaining end-to-end latency of 31ms, 32 FPS: 8ms voxelization, 19ms inference, 4ms NMS post-processing.
C++TensorRTCUDAPOSIX ThreadsONNXPointPillars3D Object DetectionLiDAR VoxelizationMultithreading
VectorNet – Graph-Based Trajectory Forecasting
Jan 2025 – May 2025
- • Implemented VectorNet: a two-stage Graph Neural Network (GNN) for motion prediction in autonomous driving using the Argoverse 1.0 dataset, involving both local polyline-based SubGraphs and global self-attention modules.
- • Performed hand-crafted feature extraction from raw CSVs and constructed per-scenario graph datasets for ego agent, lane, and surrounding objects; visualized and evaluated trajectory predictions against ground truth.
PyTorch GeometricGNNGraph AttentionArgoversePyTorchPickleNetworkXCondaJupyterPython
BioMedRAG - Real-Time Retrieval and Generation for Biomedical Data
Jan 2025 – Feb 2025
- • Engineered Retrieval-Augmented Generation (RAG) system integrating BioGPT – a finetuned large language model with NCBI Entrez - data hub for biomedical question answering, achieving an accuracy of 78% on PubMedQA dataset a 2% improvement over the state-of-the-art BioGPT—alongside a peak batch accuracy of 81.5% during evaluation.
RAGLLMGradioTransformersHugging FaceCUDATensorflowNLTKscikit-learnsacremosesPython
Object Detection for Autonomous Driving
Sep–Oct 2025
Trained Faster R-CNN and YOLOv8 on Waymo and mixed Waymo+KITTI datasets with unified 3-class taxonomy; improved cross-dataset consistency and mAP.
- • Faster R-CNN: tighter boxes + better occlusion handling.
- • YOLOv8: ~3× faster inference with strong accuracy.
CUDAYOLOv8Faster R-CNNWaymoKITTI
Fingentic — Agentic Financial AI Assistant
May–Jun 2025
Agentic platform for non-technical users to explore company financials, stock prices & trends via text + voice, delivering insights with charts, tables, and cited summaries.
- • LangChain + LangGraph tool orchestration, SQL generation, web search fallback, guardrails & transparent citations.
- • Realtime APIs + voice workflow (Whisper) for conversational finance UX.
LangGraphGemma-27BSQLiteGradioyFinanceGuardrails
Capstone project: Skin Cancer AI, Thapar Institute of Engineering and Technology
Jan 2022 – Dec 2022
- • Built a portable real-time AI Device to distinguish between benign and malignant Skin Cancers by Finetuning MobileNet V2 model achieving 94.2% accuracy on the final prototype orchestrated using Intel® Movidius™ Neural Compute Stick, Raspberry Pi 3b+ and a camera module for early detection of skin cancer in remote areas with limited healthcare access.
TensorFlowKerasPandasOpenCVNumPyCNNTransfer LearningFrozen GraphAI on the Edge