
Commerce · End-to-end
Retail customer analytics
From 541,909 invoice lines to a clear view of sales concentration, customer segments, and repeat purchasing.
DATA SCIENCE & ANALYTICS PORTFOLIO
From complex data
to decisions that make sense.
Exploring real-world questions through Python, SQL, statistics, and machine learning. Every project starts with a question and ends with evidence.
Discover my worklower forecast error than the
previous-week baseline

01 / SELECTED WORK
Four studies. Different questions.
The same commitment to evidence.

Commerce · End-to-end
From 541,909 invoice lines to a clear view of sales concentration, customer segments, and repeat purchasing.

Mobility · End-to-end
A next-day forecast with calendar-aware features and an honest test on future observations.

Materials · Modeling
Predicting concrete strength while keeping repeated mix formulations on the same side of the evaluation split.

Manufacturing · Modeling
Recognizing steel plate fault types, with target leakage removed and class-level performance made visible.
NEW / POWER BI PROJECTS
Editable reports, data models, and DAX.
Built around real analytical questions.
Report formats and prepared data are checked. Native Power BI Desktop refresh and rendering remain to be verified. The figures below show the Python analysis.

POWER BI · Commerce
Three editable report pages, 17 DAX measures, prepared data, and an executed analysis notebook.

POWER BI · Mobility
Three editable report pages, 17 DAX measures, prepared data, and an executed analysis notebook.
02 / THE PROJECT LIBRARY
Explore by dataset, question, or method.
Every study has a result you can inspect.
Two end-to-end case studies, 98 focused analyses, and two advanced Power BI projects across 13 historical public datasets. Related studies reuse data to explore different questions.
102 projects
Where are invoiced sales concentrated, and which customers should be considered for a retention experiment?
Can historical demand and calendar information improve next-day hourly rental forecasts?
Which product codes have the largest credited value relative to invoiced sales?
Which frequently purchased product pairs co-occur more than their individual popularity suggests?
How many product codes account for the first 80% of invoiced product value?
How do plausible cleaning choices alter gross invoiced sales?
What is the distribution of elapsed time between identified customers’ purchases?
How does product variety per invoice relate to order value and units to handle?
Which frequently sold codes exhibit the widest middle-range price differences?
How does repeat purchasing compare when every eligible customer receives 90 days of follow-up?
How does the registered-rider share differ by day type and selected hours?
How does observed commute-hour demand vary across weather categories and day types?
When are calendar hours missing from the rental dataset?
At which elapsed-time lags are observed hourly rental counts most strongly correlated?
What share of a complete day’s rentals occurs in its four busiest hours?
How does observed hourly demand compare between corresponding months of the two years?
How do observed subscription rates differ by recorded contact channel?
How does observed subscription vary with the number of campaign contacts?
Do prior successful campaign records identify a different response population?
Can contact context and prior campaign history predict subscription without call duration?
How much does unavailable call-duration information change retrospective model scores?
Do predicted probability bands match observed subscriptions on later records?
What precision and recall would observed rankings yield when reviewing a fixed fraction of records?
Which fields contain the most explicit unknown values, and does their response mix differ?
How does response differ between never-contacted clients and prior-contact recency bands?
How do response and campaign intensity change across source-order quarters?
How strongly are sensory labels concentrated in the middle of the rating scale?
How does median sensory quality vary across declared alcohol bands within red and white samples?
Which laboratory measurements have the strongest rank associations with quality within each wine type?
How does density vary across residual-sugar bands within wine type?
How variable is the free-to-total sulfur measurement ratio?
Can chemical measurements estimate sensory quality better than a median baseline?
Can chemistry rank samples with an observed quality score of at least seven?
How often do identical laboratory signatures recur, and can their sensory scores disagree?
Does a pooled quality model have different error or bias for red and white samples?
Do chemistry-based clusters differ in their sensory-score distributions?
How do strength distributions differ across curing-age bands?
How does observed water-to-binder ratio relate to strength?
How do observed supplementary-binder fractions compare in 28-day samples?
How does water content differ across observed superplasticizer dosage bands?
Can mix measurements predict strength beyond a median baseline on unseen formulations?
At which curing ages are the fixed strength model errors largest?
Which input variables carry overlapping rank information?
Can mix and age measurements rank specimens below a declared analytical strength threshold?
Which hours have the highest typical appliance energy per recorded interval?
How do interval energy distributions differ across weekdays?
How variable is daily appliance use when only complete days are counted?
What fraction of appliance energy occurs in the highest-load recorded intervals?
Which recorded indoor conditions have the strongest rank associations with appliance energy?
Can lagged energy and known calendar features estimate appliance use an hour ahead?
How do source-supplied random controls compare with measured variables in an association screen?
Can past appliance use rank unusually high-load intervals one hour ahead?
Which sensor and reference fields have the weakest observation coverage?
How long are the longest consecutive missing runs in each monitoring channel?
What daily pattern appears in observed reference carbon monoxide measurements?
Which sensor responses are most associated with each available gas reference?
Can contemporaneous sensor responses estimate an available reference CO concentration?
How stable is the relationship between a CO-sensitive sensor and the reference over calendar months?
Can earlier observed CO and sensor responses improve a one-hour-ahead reference forecast?
Does the fixed CO calibration model perform differently across relative-humidity bands?
How does observed purchase conversion vary across recorded months?
Do recorded visitor types differ in purchase conversion and product engagement?
Which recorded traffic categories combine meaningful volume with higher observed conversion?
How does completed-session product browsing depth relate to conversion?
How do bounce and exit rates relate, and where are their recorded values unusual?
How well can a model classify purchase sessions without PageValues?
How does PageValues alter retrospective purchase classification?
Which browser categories warrant investigation of their purchase experience?
How does observed city-cycle fuel economy vary by model year?
How does fuel economy differ across declared vehicle-weight bands?
How do cylinder groups differ in observed fuel economy and vehicle weight?
Does the origin comparison change when model years are grouped into comparable periods?
Can technical specifications estimate historical city-cycle fuel economy?
How do specification-based models perform on later model years?
Which model-year and origin groups contain missing horsepower measurements?
How do MPG summaries compare with fuel used per 100 kilometres?
How does the median spending mix differ between recorded sales channels?
What fraction of aggregate annual spending is associated with the highest-spending customers?
Do customers form interpretable groups based on annual spending patterns?
Which spending categories show the strongest customer-level rank associations?
How does observed total annual spending differ across recorded regions?
Which customer spending profiles are unusual relative to this sample?
Which defect classes dominate the labeled inspection sample?
How do pixel area and perimeter profiles differ by recorded fault class?
How does the composition of recorded faults vary across plate-thickness bands?
Which fault classes have different observed image contrast and luminosity profiles?
Can inspection features distinguish the seven recorded fault classes?
Which true fault classes are most often confused with another class?
How much recognition performance changes when only compact geometric features are retained?
Do image measurements contain impossible bounds or nonpositive areas?
Which months contain the most fire records and recorded burned area?
How much of the recorded burned area comes from the largest observations?
How does the share of zero recorded area vary by month?
Which recorded weather and fire-index variables associate with burned area?
Which source grid cells contain the greatest recorded burned area?
Can recorded conditions estimate burned area beyond a simple baseline?
Can physical measurements estimate observed ring counts beyond a median baseline?
How does restricting predictors to external measurements change ring-count error?
Which observations violate simple dimensional or mass consistency checks?
How do shell-shape ratios and observed rings differ among recorded specimen categories?
How can a commercial team monitor sales, credits, and customer coverage without ambiguous KPI definitions?
Where does a better next-day demand forecast still underpredict, and how much of the evaluation calendar is observed?
Try a different keyword or dataset.
03 / MY APPROACH
I use data science to connect a real question with an answer that can be checked. My work spans customer behavior, forecasting, quality analysis, and model evaluation—with the assumptions in view.
Define the decision, the available information, and what a useful answer would change.
Trace sources, preserve missingness, prevent leakage, and compare with a simple baseline.
Share code and findings alongside uncertainty, limitations, and the next test worth running.
BUILT TO BE EXPLORED
Browse the source, open an executed notebook, or reproduce an analysis.
The complete collection includes setup instructions and dataset attribution.