Hey, I'm

Tushar

Data Scientist · Machine Learning Engineer

Top 10 of 60+ teams, Penn State Nittany AI Challenge 2026
92% forecasting accuracy on historical attendance data at Penn State
305K+ tweets analyzed in a global LLM sentiment pipeline
2,277 job descriptions semantically ranked via deep learning
10M+ billing records validated across telecom pipelines at Oracle

$ git log --oneline --graph

How I got here

Five years of turning messy data into statistical evidence set the direction. Every stop since has been about building models and systems that hold up under real scrutiny, not just look good in a notebook.

7f3a2c1 Accenture · Jun 2019 – Feb 2022

feat: SQL-based data validation framework, -60% discrepancies

First real lesson in production data: a 4.5M-customer billing base doesn't forgive weak validation. Designed row-level checks and reconciliation logic that cut discrepancies 60% and shortened month-end close by 10%.

c9e81b4 Oracle Corporation · Feb 2022 – Dec 2024

feat: anomaly detection at scale, 600K+ subscribers

Moved from fixing issues after the fact to designing the checks that catch them first — anomaly detection and reconciliation holding 98%+ accuracy within SLA. Earned "Best Upcoming Talent," FY23 Q3.

a10d5e7 Penn State Research · Sep 2025 – May 2026

feat: forecasting + agentic auditing, 92% accuracy

Went back to build the theory behind the practice: statistical forecasting models and a first prototype of autonomous agents doing continuous compliance auditing.

f42b901 Barton Malow · May 2026 – Aug 2026

feat: AI cost-attribution analysis, 5.5x growth surfaced

Where the two threads met: quantitative analysis of AI agent costs inside shared compute, and a 211-rule data quality framework shipped to a $159B production pipeline.

The throughline: every role added statistical rigor, first to validation, then to detection, then to forecasting, then to AI-driven analysis. The Barton Malow work is what all four look like combined into one.

$ cat PRINCIPLES.md

The pragmatic builder

Everything I build is grounded in a production mindset.

"Understanding existing systems deeply before changing them, designing for the next person who has to maintain my work, and communicating technical findings in a way non-technical stakeholders can act on." That’s the standard I actually build to.

$ cat experience.log

5+ years turning messy real-world data into models and evidence leadership can act on

Data Science Intern

Barton Malow Featured

May 2026 – Aug 2026 Michigan, United States
  • Applied quantitative cost-attribution analysis and an automated ownership-resolution hierarchy across 933 resources on Databricks, surfacing $142K in lifetime platform spend and $13.2K in idle resource cost across 885 unused objects, and cutting the unowned-spend gap from 85.6% to 54%.
  • Built a Claude AI cost-attribution dashboard isolating agent query costs from shared Databricks serverless compute (20+ users, 8,000+ queries, $0.53 avg cost/query), revealing AI's share of warehouse spend grew 5.5× in 5 months (1.5% → 8.3%), giving leadership real-time cost visibility before it became a budget concern.
  • Built a 211-rule data quality framework spanning 23 production tables in a $159B CRM pipeline, achieving 100% clean execution while using an AI-driven profiler to accelerate rule generation.
PythonQuantitative AnalysisDatabricksPySparkSQLClaude Code

Research Assistant, Data Science

Pennsylvania State University

Sep 2025 – May 2026 Pennsylvania, United States
  • Developed statistical forecasting models on historical attendance data, achieving 92% accuracy and enabling data-driven vendor allocation and capacity planning.
  • Reduced event planning timeline 83% (12 months → 2 months) by building data-driven workflows, dashboards, and scenario simulations that aligned stakeholders on constraints and priorities.
  • Prototyped autonomous AI agents for continuous auditing and ethical compliance, including WORM-based audit trails, ontology-driven fairness checks, and cybersecurity safeguards aligned with ISO/IEC 42001.
Statistical ModelingForecastingAgentic AIPythonSQL

Staff Data Consultant

Oracle Corporation

Feb 2022 – Dec 2024 Bangalore, India
  • Applied SQL-driven anomaly-detection and reconciliation analysis across 600K+ telecom subscriber billing records, maintaining 98%+ accuracy within SLAs.
  • Built and automated KPI dashboards (churn, usage patterns) with 99% on-time delivery, enabling leadership to track performance in weekly reviews.
  • Led onsite test efforts, coordinating with offshore analytics teams to deliver readiness dashboards supporting a successful, on-time go-live.
SQLData AnalysisOracle DatabaseBusiness Analysis
Award Oracle "Best Upcoming Talent," FY23 Q3

Data Consultant

Accenture Solutions Pvt Ltd

Jun 2019 – Feb 2022 Pune, India
  • Identified and resolved 50+ recurring data quality issues in production billing datasets for a 4.5M-customer base, improving billing accuracy and customer trust.
  • Automated batch bill-run workflows using SQL scripts and scheduling, cutting processing time from 1 day to 2 hours (90× faster) and manual effort by 92%.
  • Designed a SQL-based data validation framework (row-level checks, reconciliation, exception flags) that reduced discrepancies by 60% and shortened month-end close by 10%.
SQLShell ScriptingExcelClient Communication
Award Accenture "Shared Success Catalyst," Addressing Client & Community Needs

$ ls ./projects

Real systems, real numbers

5CMS rules mapped
Top 10of 60+ teams
100%deterministic fallback

AI Sentinel: Compliance Monitoring

Autonomous compliance agents detecting AI-driven regulatory risk in real time. Top 10 of 60+ teams, Penn State Nittany AI Challenge 2026. Hybrid architecture: deterministic rule engine + local LLM explanation layer with a 100%-deterministic fallback.

Every finding SHA-256 audit-logged Deterministic fallback: no unchecked LLM decisions
PythonStreamlitOllamaGemini
View repo →
120resumes ranked
2,277job descriptions
384-dSBERT embeddings

Resume–JD Matching via Deep Learning

3-stage semantic ranking pipeline improving over keyword-based ATS matching: TF-IDF baseline → SBERT semantic embeddings → supervised neural refinement into a calibrated match probability.

Baseline before complexity: TF-IDF → SBERT → neural Calibrated probability output, not a raw score
PythonPyTorchSBERTscikit-learn
View repo →

Worldwide LLM Sentiment Analysis

End-to-end NLP + geospatial pipeline tracking global sentiment toward LLMs across 305K+ tweets. Location-normalization cut unresolved geography from 72.8% to 43.85%.

PythonNLTK/VADERPlotly
View repo →

Deep Dive · Penn State Nittany AI Challenge, 2026

AI Sentinel: Automated Compliance Monitoring

Problem

Fall-monitoring AI systems can quietly restrict resident movement in ways that violate federal nursing-home regulations, and nobody was watching for it in real time.

Why

A missed compliance risk isn't just a fine, it's a resident's autonomy on the line. And a black-box LLM verdict isn't good enough when the finding has to hold up to an auditor.

How

A hybrid architecture: a deterministic rule engine maps directly to CMS regulations, a local LLM explains each finding in plain language, and a 100%-deterministic fallback means the LLM never gets the final say.

Findings covered by the deterministic fallback

100%
5CMS regulations mapped (42 CFR §483)
Top 10of 60+ teams, Nittany AI Challenge
100%deterministic fallback coverage
2funded phases won (Prototype + MVP)

Deep Dive · Independent Project, 2026

Resume–JD Matching via Deep Learning

Problem

Keyword-based ATS matching misses qualified candidates who describe their skills differently than a job description does, and over-ranks resumes that just repeat the right buzzwords.

Why

A ranking system that can't tell semantic fit from keyword-stuffing produces bad shortlists on both ends: real candidates filtered out, weak ones let through.

How

A 3-stage pipeline that only adds complexity when it earns its place: a TF-IDF baseline first, then SBERT semantic embeddings, then a supervised neural layer that outputs a calibrated match probability instead of a raw similarity score.

120resumes ranked
2,277job descriptions
3pipeline stages, each measured before the next
384-ddense embeddings (all-MiniLM-L6-v2)

$ cat research.md

Published work and competitive recognition

Published Paper · Springer Nature

Autonomous Multi-Agent Governance: AI Sentinel Framework for Mitigating Psychological Restraints in AI-Driven Fall Management

Accepted and presented at SEET 2026 (2nd International Conference on Software Engineering of Emerging Technologies), Penn State Behrend, PhD Research Track. Listed on page 76 of the conference abstract book. To be published in the official conference proceedings by Springer Nature.

View conference →
Competition · Penn State Nittany AI Challenge 2026

AI Sentinel: Top 10 of 60+ Teams

Won the Prototype Phase and MVP Phase, advanced through funded phases to a Top 10 finish at the Pitch Contest, Hintz Family Alumni Center, University Park.

Award · Oracle Corporation

"Best Upcoming Talent" (FY23 Q3)

Recognized for exceptional contributions and high-impact performance on the Telia Company billing analytics engagement.

Leadership · Penn State University

Student Senator

Representing students in discussions with university leadership; championing AI literacy and equitable AI access across Penn State's Commonwealth Caucus.

$ cat skills.yaml

Tools I actually ship with

ML & Statistics

Regression · Classification · Clustering · NLP · Time Series Forecasting · Feature Engineering · Model Evaluation (AUC, F1)

Programming

Python (Pandas, NumPy, Matplotlib, Seaborn, scikit-learn, TensorFlow, PyTorch, Keras) · SQL · R · Java

AI Tools

Claude Code · Claude Cowork · Claude Skills · Claude Projects · Claude MCP · ChatGPT · GitHub Copilot · Cursor IDE · Perplexity · Ollama · Google Gemini · Julius AI

BI & Visualization

Tableau · Power BI · Plotly · Matplotlib · Seaborn

Data Engineering & Cloud

Databricks · PySpark · Unity Catalog · Azure · AWS · Oracle Database · PostgreSQL · PL/SQL

Dev Tools

Git · GitHub Actions · Jupyter · Jira

$ cat education.log

Background

Pennsylvania State University

Aug 2025 – Dec 2026

Master of Data Analytics · GPA 3.96/4

Statistical Analysis, Data-Driven Decision Making, Predictive Analytics, Deep Learning, Natural Language Processing.

KIIT University

May 2015 – May 2019

B.Tech, Electronics & Instrumentation Engineering · GPA 3.3/4

Control Systems, Digital Electronics, OOP, Data Structures & Algorithms, Web Technology, Artificial Intelligence.

$ mail tushar

Let's build something that ships.

Open to full-time Data Scientist / Machine Learning Engineer roles starting Dec 2026.