Building Analytics Infrastructure | 5+ Years Experience | Gen AI Ready
This portfolio demonstrates my automation-first approach to building BI systems across three distinct contexts:
Scaling Enterprise Infrastructure - P&G Supply Chain Analytics
Building from Zero for Startups - Via.work BI Platform
Leveraging AI to Eliminate Manual Processes - GenAI-Powered Classification
Each project follows a consistent framework: Problem β Architecture β (Automation as needed) β Outcome
Core Challenge: Modernise a fragmented, 20+ day manual reporting process into an automated, scalable analytics platform serving 24 APAC markets and data across 28,000 stores
Business Context
P&G's legacy distribution tracking system was unable to support proactive decision-making:
Data scattered across 6+ disparate systems (Azure, SQL, SAP, manual exports)
Dashboard refresh took 20+ days, making insights obsolete before delivery
No unified view across 24 countries, 20,000+ stores, 60M+ rows of store-SKU data
100+ Country and Category managers lacked self-service capabilities
Technical Debt
Siloed data schemas with no harmonization
Manual data pulls and transformations
No security model for multi-market access
Poor performance at scale
Stakeholder Impact
Business and IT Country Heads couldn't make timely supply chain optimization decisions, resulting in missed opportunities for distribution improvements and inventory management.
Approach: Start small, prove value, earn trust
βββββββββββββββββββ
β Single Country β β Manual Data Pulls β PowerBI Desktop β Category Manager
β (Thailand) β (Data Model + Local Testing)
βββββββββββββββββββ
Key Activities:
Worked directly with Singapore Category Manager to understand pain points
Built prototype dashboard with core Distribution KPIs
Demonstrated value: Reduced their reporting time from 5 days to 2 hours
Won stakeholder trust through iterative feedback and quick wins
Technologies Used:
Power Query (data cleaning & transformation)
DAX (custom measures and KPIs)
Power BI Desktop (visualization)
Outcome:
Secured buy-in from Country Head to scale solution across entire APAC region
Approach: Architect for scale, automate everything, ensure quality
βββββββββββββββ ββββββββββββββββ βββββββββββββββ ββββββββββββββββ
β Data Sourcesβ β β ETL Layer β β β Semantic β β β Power BI β
β (6 schemas) β β β β Layer β β Premium β
β β β β β (Azure AAS)β β (P2 Tier)β
βββββββββββββββ ββββββββββββββββ βββββββββββββββ ββββββββββββββββ
β β β β
ββ Azure DB ββ Harmonization ββ DAX Measures ββ RLS (Country)
ββ SQL Server ββ Data Quality ββ Business Logic ββ Self-Service
ββ SAP ββ Transformation ββ Semantic Model ββ Incremental
ββ Manual Exports ββ Single Source ββ Refresh
Team Structure:
Me (Lead Consultant): Solution architecture, stakeholder management, semantic layer design, Power BI development + 2 BI Developers
2 External Data Engineers: ETL pipeline development, data quality frameworks, SQL ETL experts
Total Team: 5 people
Architectural Decisions:
Data Engineering Layer (Azure)
Automated data ingestion from 6 disparate sources
Built data quality validation framework
Implemented data harmonization logic
Created staging β bronze β silver β gold data pipeline
Semantic Layer (Azure Analysis Services)
Moved complex calculations from Power BI to backend
Designed star schema with fact and dimension tables
Centralized business logic for consistency
Optimized for query performance
Consumption Layer (Power BI Premium P2)
Deployed to enterprise-grade infrastructure
Implemented Row-Level Security (RLS) at country level
Used bookmarks for user-centric page navigation
Optimized visuals for 100+ concurrent users
Refresh Schedule:
Incremental Refresh (Daily): New/updated data only, refreshes overnight (~4 hours)
Full Refresh (Weekly): Entire dataset, runs over weekend (~18 hours)
What We Automated:
Data Ingestion
Eliminated manual exports from SAP, SQL, Azure
Built scheduled Azure jobs for automatic data pulls
Implemented retry logic and error handling
Data Quality Checks
Automated validation rules for data completeness
Anomaly detection for unusual data patterns
Alerting system for data quality issues
Transformation & Harmonization
Standardized 6 different schemas into unified model
Automated currency conversions and unit standardizations
Built reusable transformation templates
Dashboard Refresh
Moved from manual β scheduled incremental refresh
Implemented intelligent caching strategies
Optimized refresh for performance (incremental vs. full)
Distribution & Access
Self-service model: Users access dashboards directly
Automated RLS provisioning based on user roles
Eliminated manual report distribution
Why Automation Was Critical:
Scale: 24 markets Γ multiple categories = impossible to maintain manually
Timeliness: Business decisions needed current data, not 20-day-old data
Consistency: Automated pipelines ensure standardized calculations across markets
Reliability: Reduced human error and dependency on individual analysts

Business Impact:
Metric | Before | After | Impact |
|---|---|---|---|
Dashboard Refresh Time | 20+ days | 24 hours | 95% reduction |
User Access | 5 analysts + 1 Stakeholder | 100+ stakeholders | 20x scale |
Data Freshness | Bi-weekly | Daily (incremental) | 30x improvement |
Manual Effort | 120 hours/month | 0 hours/month | 100% elimination |
Markets Covered | 1 (Singapore pilot) | 24 APAC markets | 24x expansion |
Specific Use Cases Enabled:
Out-of-Stock Prevention: Identified distribution gaps before they became stockouts
Retail Execution: Tracked in-store visibility and shelf placement effectiveness
Promo Performance: Measured real-time impact of marketing initiatives
Market Comparison: Benchmarked performance across countries
Financial Impact
Project success led to $50,000 additional revenue for my consulting firm through contract expansion
Start Small, Scale Fast: Phase 1 MVP earned stakeholder trust, making Phase 2 scaling possible
Move Compute to Backend: Shifting calculations from Power BI to semantic layer dramatically improved performance
Design for Self-Service: Empowering users to explore data reduces ad-hoc requests and scales better
Automation Requires Investment: Building robust pipelines takes time upfront but pays massive dividends at scale
Data Quality is Non-Negotiable: Automated quality checks maintain stakeholder trust in self-service environment
Core Challenge: Build automated financial reporting and cross-functional analytics infrastructure during critical acquisition phase
Business Context:
Via.work, a Series B HR/payroll startup, was undergoing acquisition by Justworks:
CEOs needed automated monthly financial updates for investors and board
No unified view across Sales, Operations, and Finance verticals
Manual consolidation of CRM data, revenue metrics, and operational KPIs
Teams working in silos with no cross-functional visibility
Technical Gap
No existing BI infrastructure or analytics team
Data scattered across Hubspot CRM, Stripe (payments), spreadsheets
Manual monthly report generation taking 3-4 days
High risk of errors in critical acquisition-period reporting
Stakeholder Impact
CEOs spending valuable time consolidating data manually instead of focusing on acquisition strategy; teams unable to understand how their actions impacted other verticals or company bottom line.
Approach: Build minimum viable analytics infrastructure that grows with the company
ββββββββββββββββββββ βββββββββββββββββββ ββββββββββββββββββββ
β Data Sources β β β Transformation β β β Consumption β
β β β & Semantic β β Layer β
β β β Layer β β β
ββββββββββββββββββββ βββββββββββββββββββ ββββββββββββββββββββ
β β β
ββ Hubspot CRM ββ Power Query ββ Executive
β (Sales data) β (ETL logic) β Dashboard
β β β (Investor
ββ Stripe + Internal SQL ββ DAX Semantic β Updates)
β (Operational) β Layer β
β β (KPIs & Metrics) ββ Sales
ββ Google Sheets β β Dashboard
(Revenue Data) ββ Data Validation β (Pipeline)
β Rules β
β ββ Operations
ββ Automated Refresh Dashboard
Schedule (Efficiency)
Built from Scratch: Since no BI infrastructure existed, I architected the complete solution:
Data Integration Layer
Connected to Hubspot CRM via Python Pulls
Integrated SQL operations data
Consolidated operational data from Google Sheets and internal systems
Semantic Layer (Power BI)
Designed data model connecting all sources
Created business logic for key metrics:
Financial: MRR (Monthly Recurring Revenue), ARR, Churn Rate, ARPU (Average revenue per Unit)
Sales: Pipeline velocity, Lead-to-Customer conversion, CAC (Customer Acquisition Cost)
Operations: Customer onboarding time, Support ticket resolution
Built calculated measures translating technical data into business KPIs
Dashboard Layer
Executive Dashboard: Monthly board updates with financial health metrics
Sales Dashboard: Pipeline tracking, conversion funnels, win rates
Operations Dashboard: Operational efficiency, customer success metrics
Cross-Functional View: How each vertical impacts company performance
What We Automated:
Financial Reporting (Built on Google Sheets
Data Validation
Cross-Functional Insights
Report Distribution
Why Automation Was Critical:
Acquisition Timing: Every day counted; CEOs needed to focus on deal, not data
Accuracy: Manual consolidation risked errors in critical investor communications
Agility: Business moving fast; needed real-time insights to adapt strategy
Scalability: Built foundation for post-acquisition analytics needs



Decision-Making Transformation
Before:
Teams operated in silos without visibility
Financial reporting was reactive and manual
No understanding of cross-functional impact
After:
Teams understood how their work impacted company goals
Proactive identification of trends and issues
Data-driven conversations across all verticals
Reliable, automated financial reporting built investor confidence
Clean data room accelerated due diligence process
Specific Use Cases Enabled:
Investor Communication: Real-time financial health metrics for board meetings
Sales Strategy: Pipeline visibility enabled better forecasting and Operations Team resource allocation
Strategic Planning: Cross-functional data informed product and go-to-market decisions
Data Sources:
Hubspot CRM (API integration)
SQL (client data)
Google Sheets (operational + Revenue data)
Internal databases
ETL & Transformation:
Power Query (data transformation)
Custom API connectors + Python
Semantic Layer:
Power BI data modeling
DAX (calculated measures and KPIs)
Relationship management
Visualization:
Power BI Desktop (development)
Power BI Service (deployment)
Shared workspaces for collaboration
Key Features:
Self-service dashboards
Cross-source data validation
Mobile-responsive design
Build for Self-Service Early: Empowering teams to access their own data reduced my bottleneck as solo BI person
Start with CEO Pain Points: Solving executive reporting needs first secured buy-in for broader analytics investment
Semantic Layer is Critical: Centralizing business logic in data model ensured consistency across all dashboards
Automation Builds Trust: Reliable, automated reports during acquisition built confidence with investors
Cross-Functional Visibility Drives Collaboration: When teams see how they impact each other, they work together better
Core Challenge: Automate manual 3-language job classification while ensuring GDPR compliance and mitigating algorithmic bias
Business Context
Luxembourg's labor market operates in three official languages (French, German, English):
Government needed to classify job advertisements into standardized skill categories
Manual process required analysts to read and categorize each job ad
Language barriers made consistent classification difficult and biased
Scale problem: Thousands of job ads monthly, growing backlog
Technical Challenges
Multilingual NLP: Models trained on one language perform poorly on others
GDPR Compliance: Cannot send sensitive job data to external LLM APIs
Algorithmic Bias: Risk of model favoring certain languages or demographic groups
Data Scarcity: Limited labeled training data across all three languages
Stakeholder Impact
Labor market analysts spending weeks manually classifying jobs instead of analyzing trends; inconsistent categorization across languages undermining data quality for policy decisions.
Approach: Build GDPR-compliant, bias-aware AI pipeline for automated classification
βββββββββββββββββββ ββββββββββββββββββββ βββββββββββββββββββ ββββββββββββββββ
β Job Ads β β β Translation β β β Classification β β β Validation β
β (FR/DE/EN) β β Layer β β Layer β β & Output β
β β β (Self-hosted β β (mBERT) β β β
βββββββββββββββββββ β LLM) β β β ββββββββββββββββ
β ββββββββββββββββββββ βββββββββββββββββββ β
β β β β
ββ French ads ββ Offline hosting ββ Multilingual ββ Bias
ββ German ads β (GDPR compliant) β BERT model β metrics
ββ English ads β β β
ββ GenAI-powered ββ Fine-tuned on ββ Skill
β translation β job data categories
β β
ββ Data stays on-prem ββ Synthetic data
for bias reduction
System Components:
Translation Layer (Self-Hosted LLM)
Deployed open-source LLM locally for GDPR compliance
Used for translating job ads to common language
Offline workflow: No data leaves government infrastructure
GenAI-powered semantic understanding beyond word-for-word translation
Classification Layer (mBERT)
Multilingual BERT model (mBERT) as core classifier
Fine-tuned on labeled job advertisement data
Trained to classify into standardized skill taxonomy
Bias Mitigation Framework
Generated synthetic training data using LLMs
Balanced representation across languages
Statistical testing for fairness across demographic groups
Monitoring dashboard for ongoing bias detection
Validation & Quality Assurance
Human-in-the-loop validation for edge cases
Confidence scoring for each classification
Feedback mechanism for continuous improvement
Why Automation + AI Was Critical:
Scale: Manual process couldn't keep pace with job ad volume
Consistency: Eliminated subjective interpretation differences between analysts
Language Equality: Ensured no language disadvantaged in classification
Efficiency: Freed analysts to focus on policy analysis vs. manual categorization
Privacy: Self-hosted solution maintained GDPR compliance
AI/GenAI Usage:
Use Case | Technology | Purpose |
|---|---|---|
Translation | Self-hosted open-source LLM | Convert job ads to common language while staying GDPR-compliant |
Classification | mBERT (Multilingual BERT) | Classify job ads into skill categories across three languages |
Synthetic Data | Fine-tuned LLM | Generate balanced training data to reduce bias |
Bias Analysis | Statistical testing + monitoring | Ensure fairness across languages and demographics |
Business Impact
Metric | Before | After | Impact |
|---|---|---|---|
Classification Time | 100-200 ads/day (manual/rule based) | 1000+ ads/hour | 100x faster |
Consistency | Variable (analyst-dependent) | Standardized (model-based) | Objective classification |
Language Coverage | Required trilingual analysts | Works across FR/DE/EN equally | Expanded capacity |
Bias Detection | No systematic approach | Automated monitoring | New capability |
GDPR Compliance | Manual = compliant | Automated + compliant | Privacy maintained |
Research Contributions
GDPR-Compliant AI Architecture
Demonstrated how to deploy GenAI/LLMs offline for sensitive government data
Provided framework for EU organizations to leverage AI while maintaining compliance
Advisory report submitted to Luxembourg and Dutch governments
Multilingual Bias Mitigation
Analyzed algorithmic bias across three languages
Developed methodology for ensuring language equality in NLP systems
Published techniques for synthetic data generation to reduce bias
Production-Ready POC
Delivered working proof-of-concept system
Documented deployment architecture for government IT teams
Created monitoring dashboard for ongoing quality assurance
Specific Use Cases Enabled:
Labor Market Analysis: Faster, more consistent skill demand tracking
Policy Making: Reliable data on skill gaps for education policy
Job Seeker Support: Better matching between job ads and training programs
Data & Infrastructure:
Python (data processing & model training)
Self-hosted LLM (translation)
On-premise compute (GDPR compliance)
Machine Learning:
mBERT (Multilingual BERT) - core classifier
Hugging Face Transformers library
PyTorch (model training framework)
Scikit-learn (evaluation metrics)
GenAI/LLM Components:
Open-source LLM (offline translation)
Fine-tuned LLM for synthetic data generation
Prompt engineering for data augmentation
Bias & Quality:
Statistical significance testing
Fairness metrics (demographic parity, equal opportunity)
Custom monitoring dashboard
Human-in-the-loop validation interface
Key Innovations:
GDPR-compliant LLM hosting architecture
Synthetic data generation for bias reduction
Multilingual model evaluation framework
Production deployment guidelines for government use
GDPR and AI Can Coexist: Self-hosted LLMs enable GenAI benefits while maintaining compliance
Bias is a Feature, Not a Bug: Proactive bias detection must be built into AI systems from day one
Synthetic Data is Powerful: LLM-generated training data can reduce real-world bias when done carefully
Multilingual NLP is Hard: Language equality requires intentional design and continuous monitoring
Government AI Needs Different Approach: Privacy, explainability, and fairness take precedence over raw performance
These three projects demonstrate my core philosophy: Build analytics systems that eliminate manual work through intelligent automation.
Automation-First Mindset
P&G: 20+ days β 24 hours through automated pipelines
Via.work: 3-4 days β 30 minutes for financial reporting
GenAI Prototype: 10-20 ads/day β 1000+/hour classification
Semantic Layer Design
P&G: Azure AAS semantic layer for 60M+ rows
Via.work: Power BI data model connecting CRM + finance + ops
Self Serve Analytics as a Product
P&G: 100+ users with RLS access
Via.work: 10+ managers exploring data vs. CEO-only reports
GenAI: Analysts focus on insights vs. manual categorization
Scale Through Architecture
P&G: 1 country β 24 countries, 1 user β 100+ users
Via.work: Built for growth from acquisition through scale
GenAI: 3 languages, production-ready for thousands of job ads
Data Quality as Foundation
P&G: Automated validation rules and anomaly detection
Via.work: Cross-source reconciliation and validation
GenAI: Bias monitoring and confidence scoring
Abhiram Elangovan
π§ mail@abhiram.me
π± +31-657881547
π linkedin.com/in/elanabhi
All sensitive data has been redacted or anonymized. Happy to discuss any project in detail during the interview process.