Open to 2026 internships & early-2027 full-time roles

Data Analyst
BI Analyst

I bring 3+ years of PwC M&A analytics experience across dashboards, KPI reporting, SQL/Python validation, and executive-ready insights — now strengthened by an M.S. in Data Science at UNT (GPA 4.00) and hands-on AI/ML and LLM projects.

Targeting Data Analyst BI Analyst Analytics Engineer Data Science Intern
About

PwC analytics background, now going deeper in data science.

My strongest lane is Data Analyst and BI Analyst work — SQL, Python, dashboards, KPI frameworks, data quality, and stakeholder-facing reporting. My differentiator is graduate-level ML, NLP, and LLM project experience built on top of real consulting delivery.

I'm a data analyst with 3 years of experience at PwC's Acceleration Center in Bangalore, where I worked across Mergers & Acquisitions engagements — helping clients like GE, Pfizer, Capital One, Discover, Seagen, and NYCB navigate complex financial datasets during high-stakes deal cycles.

I'm completing my M.S. in Data Science at the University of North Texas (GPA 4.00, Dec 2026), applying that foundation through real projects in LLM evaluation, NLP, explainable AI, and multimodal modeling. I'm actively looking for Data Analyst, BI Analyst, and Data Science roles in the US where I can bring both the business context from consulting and the technical depth from graduate work.

Business analytics foundation

3 yrs PwC across M&A due diligence, financial datasets, contract analysis, KPI reporting, and executive-facing deliverables for high-stakes client engagements.

BI & data pipeline execution

SQL, Python, Power BI, Tableau, Excel, Alteryx, ETL workflows, data validation, dimensional modeling, and data-quality controls.

AI/ML project depth

Graduate work and projects covering LLM evaluation, explainable AI text detection, NLP, multimodal modeling, computer vision, and SHAP-based interpretability.

Data AnalystPrimary target
BI AnalystPrimary target
Analytics EngineerSecondary
Data Science InternSelective
Experience

3 years at PwC, now looking for the right U.S. role.

Real delivery across 6 M&A engagements: dashboards, validation pipelines, KPI frameworks, contract analytics, and executive reporting for clients like GE, Pfizer, and Capital One.

PricewaterhouseCoopers (PwC)

Data Analyst Associate 2 / Data Analyst Intern
Mar 2022 – Dec 2024 · Bangalore, India
Clients: General Electric · Pfizer · Capital One · Discover · Seagen · NYCB
  • Built and automated 7+ Power BI dashboards from KPI requirements across 6 client engagements, cutting manual reporting effort by 25%+ and accelerating executive decisions on M&A deliverables.
  • Engineered ETL pipelines using Python and SQL to validate financial data, sustaining 99.5%+ data integrity across 10+ M&A deals and reducing rework on executive deliverables.
  • Reviewed and structured 200+ legal and financial contracts per engagement using SQL-based EDA and statistical analysis, surfacing risk signals and cost-concentration patterns for client negotiations.
  • Built Python MD5 duplicate-detection automation for client data rooms, eliminating a manual process and accelerating document validation across M&A pipelines.
  • Translated ambiguous business questions into KPI frameworks and cohort analyses; partnered with PMs across 6 engagements to deliver ad-hoc visualizations and stakeholder-ready recommendations.
  • Documented data lineage and standardized finance reporting definitions across engagements using JIRA and Confluence.
SQLPythonPower BI TableauExcelAlteryx ETLKPI ReportingData Quality

University of North Texas

M.S. Data Science · Research & Projects
Jan 2025 – Dec 2026 · Denton, TX
  • Coursework: Machine Learning, NLP, Deep Learning, Data Mining, Statistical Analysis, Data Visualization.
  • Open-source contributor to SynergeReader (professor-led AI document reader) — frontend hardening, backup deployment system, port configuration, competitive analysis, and building a PDFQA-based quality-check layer to quantify feature regressions before each deployment.
  • Projects: LLM cost-performance benchmarking · Explainable AI text detection (StyleGuard) · Multimodal disaster damage assessment · AI adoption trend analysis (Tableau).
PythonPyTorchNLP LLM EvaluationPostgreSQLReact FastAPI
Projects

Projects

Each one is written around the actual problem, what I built, and what came out of it.

AI Adoption & Outcomes Analysis

ProblemHow does AI adoption relate to pay, satisfaction, and productivity across developer roles, experience levels, and org sizes?

ApproachCleaned and engineered features on ~49K Stack Overflow survey rows; built cohort segmentation, regression comparisons, and a policy-lever ROI view ranking interventions by expected uplift across role/experience/remote segments.

Outcome: executive Tableau dashboard suite showing AI adoption's strongest economic signal is amplified by foundational skill breadth (SQL, cloud, tool depth).
TableauPythonRegressionCohort AnalysisKPI Design

LLM-OpsBench: LLM Cost–Quality Benchmark

ProblemWhich LLM gives the best cost-speed-quality tradeoff for different production workloads?

ApproachConfig-driven benchmark comparing GPT, Claude Sonnet, Gemini Flash, and Llama across summarization, classification, RAG QA, and code generation. Measured quality (ROUGE-L, Macro-F1, pass@1) vs. operational cost (tokens, latency, USD). Shipped Streamlit dashboard with budget-constrained model recommender.

Outcome: no single model dominates — model selection should vary by task type, budget, and quality constraints. Demonstrated across 4 task categories with per-request telemetry.
PythonLLM EvaluationOpenAIAnthropicStreamlit

StyleGuard: Explainable AI Text Detection

ProblemExisting AI detectors give a score with no explanation — which specific writing patterns triggered it?

ApproachThree-layer pipeline: 8 stylometric features → fine-tuned RoBERTa probability scores → XGBoost ensemble with SHAP explanations for per-prediction transparency. Trained on 5K balanced arXiv abstracts (human pre-2021 vs GPT-4o mini generated).

Outcome: per-prediction SHAP explanations showing exactly which features (sentence length variance, stopword ratio, word length) drive each classification — unlike black-box detectors.
PythonRoBERTaXGBoostSHAPNLP

Multimodal Disaster Damage Assessment

ProblemHow can crisis social media and satellite imagery be fused to classify disaster damage severity in real time?

ApproachFine-tuned BERTweet for crisis-text classification (0.844 Acc, 0.795 Macro-F1) and Twitter-RoBERTa for humanitarian categorization. Implemented late fusion (text + image, α≈0.60) and trained U-Net on 256×256 satellite tiles.

Outcome: multi-modal late fusion lifted damage-severity Macro-F1 from 0.55→0.57; satellite U-Net achieved IoU 0.337 / F1 0.495 on held-out tiles. Full ETL with reproducible outputs.
PythonPyTorchBERTU-NetComputer Vision

SynergeReader: AI Document Reader (Open-source contributor)

ProblemHow do you build a reliable, grounded document Q&A system that degrades gracefully and can be validated quantitatively before each deployment?

ApproachContributed to an open-source AI document reader (React frontend, FastAPI backend, PostgreSQL + pgvector). Key contributions: frontend hardening & security, backup deployment system, port configuration, competitive landscape research, and a PDFQA-based quality-check layer to quantify answer quality before/after each release.

Outcome: introduced quantitative regression testing for document Q&A quality — replacing ad-hoc manual checks with a reproducible evaluation pipeline before each deployment.
ReactFastAPIPostgreSQLpgvectorPDFQAOpen Source
Skills

What I work with

Analytics and BI tools I use day-to-day, plus the ML and AI work I've been building through graduate projects.

Analytics Core

SQLPython / pandas / NumPy EDA & cohort analysisHypothesis testing & regression Data cleaning & validationStatistical analysis

BI & Reporting

Microsoft Power BI (DAX)Tableau KPI & dashboard developmentExcel (Pivot, VLOOKUP) AlteryxData storytelling

Data Systems

ETL/ELT workflowsData quality controls Dimensional modelingData lineage documentation PostgreSQL / pgvectorGit / JIRA / Confluence

AI / ML

scikit-learn & XGBoostNLP & transformers LLM evaluation & RAGSHAP explainability PyTorch / computer visionGenerative AI
Contact

Let's connect.

Actively looking for Data Analyst, BI Analyst, Analytics Engineer, and Data Science roles across DFW, remote, and relocation-friendly U.S. positions. Open to 2026 internships and early-2027 full-time start dates.

Role-specific resume variants (BI Analyst, Data Science) available on request.