Open to work
Plensia Lukosi · Data Analyst · Dar es Salaam, Tanzania
Building data systems.
Making data
make sense.
I build dashboards, SQL data models, and fraud-detection rules — across social content, survey data, and fintech. Three published reports below, each one starting the same way: turning a vague ask into a specific, checkable question before writing a single query.
Certified Google Advanced Data Analytics DataCamp Associate Data Analyst PL-300 · in progress
Published reports
Social content, survey data, and fintech.
LinkedIn Content Performance
“Which content is actually worth making more of?”
Organic posts from DataGirls Tanzania, Jan to Nov 2024
- Event Recap engagement (2 posts)
- 39.5%
- All-post average
- 15.1%
- Posts audited
- 59
Excel · Power QueryOpen report →
Great American Coffee Taste Analysis
“Does what people say they like match what they actually pick, blind?”
- Coffee D popularity
- 36.7%
- Coffee D blind score
- 3.38
- Respondents
- 4,023
Power BI · Power QueryOpen report →
PaySim Fraud Detection
“How do we catch fraud without freezing legitimate merchant cash flow?”
- Recall
- 97.7%
- False positives (of legitimate transactions)
- 0
- Transactions
- 6.36M
PostgreSQL · Python · Power BIOpen report →
Projects
Published reports
Each one starts the same way: turning a vague ask into a specific, checkable question before writing a single query. Pick one to open it.
LinkedIn Content Performance
“Which content is actually worth making more of?”
Organic posts from DataGirls Tanzania, Jan to Nov 2024
- Event Recap engagement (2 posts)
- 39.5%
- All-post average
- 15.1%
- Posts audited
- 59
Excel · Power QueryOpen report →
Great American Coffee Taste Analysis
“Does what people say they like match what they actually pick, blind?”
- Coffee D popularity
- 36.7%
- Coffee D blind score
- 3.38
- Respondents
- 4,023
Power BI · Power QueryOpen report →
PaySim Fraud Detection
“How do we catch fraud without freezing legitimate merchant cash flow?”
- Recall
- 97.7%
- False positives (of legitimate transactions)
- 0
- Transactions
- 6.36M
PostgreSQL · Python · Power BIOpen report →
How I read a brief
Three real briefs from my reports — and what each one turned out to actually be asking.
Report 01LinkedIn Content PerformanceOpen report →
Report 02Great American Coffee Taste AnalysisOpen report →
Report 03PaySim Fraud DetectionOpen report →
LinkedIn Content Performance
Organic posts from DataGirls Tanzania, Jan to Nov 2024. I manage the DataGirls Tanzania LinkedIn account.
The ask
"Which content is actually worth making more of?"
DataGirls Tanzania's LinkedIn presence, Jan–Nov 2024 — real organisational data, analysed as the final assignment of a data training workshop.
Definitions
The fix
// Power Query — one of three real bugs found auditing this workbook if [Content Topic] = "Program Intoduction" then "Program Introduction" else [Content Topic]
A misspelling had split one category into two. A leftover "Total" row baked into the raw table was injecting fake categories into the charts.
The finding
The posts getting the most views weren't the posts converting best.
Event Recap averaged 39.5% engagement across 2 posts, against 15.1% across all 59.
Event Recap is 2 posts (65.1% and 13.8%), so one post drives the average. Educational and Recognition & Media Coverage have 4 posts each and Event Promotion has 5. Treat these topic figures as leads to test, not proven results.
The decision
Tried a DAX column that referenced its own source field — a circular dependency. Fixing it in Power Query instead, I pasted a full step definition into a dialog that only wanted the inner expression.
Great American Coffee Taste Analysis
The ask
"Does what people say they like match what they actually pick, blind?"
Built on Maven Analytics' public Great American Coffee Taste Test dataset (4,023 survey respondents), included for the audit work, not originality.
Definitions
The fix
// one of four ordinal charts silently sorted by response count instead of sequence if [What is your age?] = "<18" then 1 else if [What is your age?] = "18-24" then 2 -- … through >65
Verified every category's canonical order against the survey's own question key before writing a single sort rule.
The finding
The coffee people said they liked best was also the one they picked blind.
4,023 respondents; four coffees blind-tasted. The stated favorite and the blind-test winner were the same coffee.
Respondents chose to take part in a public online tasting, so they are coffee enthusiasts, not a sample of all coffee drinkers. Not everyone answered every question, and the blind scores are close (Coffee D 3.38, Coffee A 3.31).
What it shows
Assumed a plain text match would sort correctly. An earlier cleanup step left trailing spaces and inconsistent capitalization, so every comparison silently failed.
PaySim Fraud Detection
The ask
"How do we catch fraud without freezing legitimate merchant cash flow?"
Self-directed case study, built on the public PaySim synthetic transaction dataset (6.36M rows).
Definitions
Try it live
Real trade-off, computed from all 6.36M transactions. Slide from $0 (flag every transfer) to $1,000,000 (flag only the largest ones) and watch recall trade against false positives.
At $200,000, this naive rule buys 66.6% recall by wrongly flagging 1,192,198 real transactions.
┊ dashed mark = where the proposed rule sits.
The finding
Full balance drain wasn't a trade-off between the two rules. It beat both, at once.
The naive fix buys recall by flagging 1.19M legitimate transactions. The proposed rule needed none of that trade.
Volume overview: 6.36M transactions, with fraud confined to CASH_OUT and TRANSFER.
Flagged accounts: the review list the proposed rule produces, filterable by day, type and amount. The .pbix file is too large for GitHub, so the three pages are shown here as screenshots.
The decision
I expected transfer velocity to be a useful fraud signal. It isn't: 99.86% of accounts transact exactly once.
PaySim is simulated data. The rule works because fraud in the simulation always empties the sender's account. In real mobile money, legitimate customers also withdraw their full balance, so this rule would need extra signals before use.
Open to work
Contacts · Dar es Salaam, Tanzania
Bring me a question you haven't finished defining yet.
Tell me the decision waiting on it, and I'll tell you what the data can actually support.
Contract work · alongside full-time roles
What I'll take on, and what you get back.
Three kinds of work I've already done and can do again on a defined scope. Each one below names the actual instance, not the capability in the abstract.
Dashboards you can maintain
The 59-post content-performance workbook in Report 01 — Power Query tagging, PivotCharts, rebuilt three times as the data kept revealing new problems. Built against a real organisation's LinkedIn export, not a clean sample file, which is where most of the actual work was.
You get A dashboard that refreshes on your own data, plus the transformation steps written down so it doesn't break the moment I stop touching it.
Data quality audit
I found a silent category-splitting typo that was misrepresenting engagement rates across a 59-post content audit, and a sort-order bug hiding inside four survey charts that everyone had read as correct. Both were invisible in the finished chart — that's the point.
You get A written list of what's wrong, what each error was doing to your numbers, and the corrected output.
SQL data modeling
Schema, views, and business-question queries in PostgreSQL against 6.36M transaction rows for Report 03, including the view that compared three candidate fraud rules on recall and false positives at once. Plus an idempotent dbt transformation layer on the Maven Fuzzy Factory pipeline.
You get Queries and views you own outright, with the assumptions behind each one documented rather than implied.
How this starts
Start with the decision, not the dataset.
Tell me the decision that's waiting on the data and roughly what shape it's in. I'll come back with what I think the work actually is — which is sometimes smaller than expected, and occasionally a different question than the one you asked.
Ready
© 2026 Plensia Lukosi ·

Tools
What I reach for, and when.
SQL — PostgreSQL & MySQL
For anything trusted twice.
Joins, aggregations, and window functions — including the fraud-rule comparison view behind Report 03's 6.36M-row dataset.
Power BI
DAX, Power Query (M), data modeling.
Rebuilt the coffee-survey dashboard after a silent sort-order bug — see Report 02's "The fix."
Excel
Power Query, PivotTables, VLOOKUP/XLOOKUP.
Built the 59-post content-performance workbook behind Report 01 — tagging, pivots, and the audit that caught a category-splitting typo.
Used inR01
Python — pandas
For verification.
Cross-checking every report's headline numbers above against the raw data before they made it into a dashboard.
Tools × projects
Where each one shows up.
Every filled cell is backed by a report on this site: solid where the tool is part of the build, light where it was used to check the numbers.
| Tool | Total | ||||
|---|---|---|---|---|---|
| SQL | Not used | Not used | In the build | In the build | 2 |
| Power BI | Not used | In the build | In the build | Not used | 2 |
| Excel | In the build | Not used | Not used | Not used | 1 |
| Python | Used to verify | Used to verify | In the build | In the build | 4 |
| dbt / Docker | Not used | Not used | Not used | In the build | 1 |
| Git / GitHub | In the build | In the build | In the build | Not used | 3 |
| Tools per project | 3 | 3 | 4 | 3 |
In the build Used to verify the numbers Click a project to open it
About
How I got here, and what I'm building next.
I'm a self-taught data analyst. I started in front-end development, which taught me to think about the end user from the start — now I run analysis and translate it into answers stakeholders can act on, the most underrated skill in this work being explaining what the numbers actually mean.
I've since built an end-to-end e-commerce analytics pipeline on the Maven Fuzzy Factory dataset — Python, PostgreSQL, and an idempotent dbt transformation layer — and I post publicly about what I'm building and figuring out, because the gap between "looks correct" and "is correct" deserves more honest documentation than it gets.
-
Front-end development Where I began — it taught me to think about the end user from the start.
-
Power BI Desktop for Business Intelligence Udemy course by Maven Analytics (Chris Dutton, Aaron Parry) — 17 hours covering DAX, Power Query, and dashboard design, Sept 17, 2025.
-
Google Advanced Data Analytics Professional Certificate via Coursera, Nov 24, 2025 — 8 courses covering statistical analysis, regression, and machine-learning fundamentals. Verify →
-
DataCamp — Associate Data Analyst Certified May 26, 2026 · DAA0013108905478 · covers data manipulation, exploratory analysis, and reporting fundamentals.
-
Maven Fuzzy Factory ETL Pipeline End-to-end e-commerce pipeline — Python, PostgreSQL, an idempotent dbt transformation layer.
-
Microsoft PL-300 Power BI Data Analyst Associate exam — September 2026.
-
Microsoft DP-600 Fabric Analytics Engineer Associate.
Certifications
Credentials, with proof.
Click any certificate to enlarge it. Microsoft PL-300 in progress — September 2026.