Open to Data Science, Analysis, AI/ML engineering and applied research roles — let's talk

Data Scientist · AI Engineer · Researcher

Hi, I'm Siddhi
I like building things with AI.

Data scientist and AI/ML researcher with 4 peer-reviewed publications.

4 Publications
3 Research venues
MS Computer Science
Curiosity

I'm a data scientist and AI enthusiast with a Master in Computer Science (Data Science specialization) from Seattle University and a Bachelor in Computer Engineering from the University of Pune. My research lives at the intersection of data quality and model behavior, the idea that better AI starts with better data, not just bigger models.

As part of a research team at Seattle University, I contributed to work on synthetic data generation for class-imbalanced medical datasets, published across IEEE Access, DaWaK, and DASFAA. Currently I'm building a personal AI memory system that captures how you think, not just what you know.

Python SQL Claude API RAG & Vector DBs LLM Applications AWS GCP Power BI Streamlit ChromaDB LangChain PyTorch Scikit-learn Flask

May 2025 — Present

Data Analyst

Boston Financial Advisory Group

Designed the data validation layer of an ML pipeline, building automated leakage checks to stop future settlement data from contaminating historical training sets, and resolved a sync-lag bug (via a SQL look-back window) responsible for 5% of apparent missing data, on top of the ETL automation that cut reporting time 40%.

Feb. 2023 — Apr. 2025

Research Data Scientist

Seattle University

ran extensive experimentation across synthetic-data generation and classification: built and benchmarked 5 SMOTE variants and 3 autoencoder architectures, designed a novel adaptive-control SMOTE algorithm, and trained/tuned multiple classical and transfer-learning classifiers on cloud GPUs — packaged into SDGnE, a reusable framework later adopted by other researchers, with results validated through a custom statistical evaluation pipeline and published across 4 peer-reviewed venues.

Jan. 2020 — Jan. 2022

Software Engineer

M.B.B. Consulting

Built cost-estimation models (Linear Regression, Random Forest) optimized on RMSE/MAE to improve pricing accuracy, and built an NLP pipeline (spaCy + regex) that cut PDF data-extraction time 80% (5 min → 1 min), saving 6.7 hours/week, alongside REST APIs and a 1M+ record ETL pipeline.

Building now

MindMirror

A personal AI system that learns how you think, not just what you know. Uses RAG and structured personal context to make every interaction feel like working with someone who has known you for months, without re-explaining yourself every time.

Claude API ChromaDB RAG Python Streamlit
Personal project

Agent Orchestration Framework

An orchestration system that wraps Claude in a plan-act-reflect loop with typed tools, human-approval checkpoints for high-risk actions, and full audit logging, turning a stateless LLM into a durable, safely-supervised autonomous agent.

Claude API FastAPI Redis Python PostgreSQL Celery Pydantic
Published · Open source

SDGnE Python Package

Open source Python package stemming from the SDGnE research project, lets users generate synthetic data from our designed algorithm for rare event and imbalanced classification tasks. Published research, usable tool.

Python Scikit-learn Synthetic Data Data-centric AI
View docs ↗
Personal project

Trail Recommendation AI Agent

An end-to-end AI agent that monitors calendar events, retrieves real-time weather data, and reasons across a personal trail database to deliver context-aware hiking recommendations, demonstrating full agent orchestration with tool use, memory, and multi-API reasoning.

n8n ChatGPT API Calendar API Weather API LLM Agents
Personal project

RAG Chatbot with Agentic Pipeline

A context-aware RAG chatbot built with LangChain and LLaMA3 fine-tuned with LoRA. Focused on production readiness — evaluating outputs critically, not just getting something that runs.

Python LangChain LLaMA3 Pinecone Streamlit
Research · Published

WalkExplorer

A cloud-hosted multimodal AI tool on GCP using CLIP transformers and OpenStreetMap data to assess urban walkability. Benchmarked against human ratings with automated test validation, published at DASFAA 2026.

GCP Python CLIP Multimodal AI OSM
Read paper ↗
Learning project

Transformer LLM from Scratch

Trained a GPT model on the Shakespeare dataset using nanoGPT with character-level tokenization and AdamW. Achieved validation loss ~1.8, built to understand transformer architecture and ML math from first principles, not just use the API.

PyTorch Python NLP Transformers

I'm currently open to AI/ML engineering, applied research, and data science roles. If you're building something interesting or just want to talk AI, I'd love to hear from you.