Data Science Group Project

I. Title slide

Predicting User Persona Segments from Ride-Hailing Behaviour

Members

01Nguyễn Giao Chi
02Trần Anh Minh
03Huỳnh Ngọc Tấn Thuận
04Huỳnh Trọng Nghĩa
05Nguyễn Quốc Khánh
06Ngô Quốc Đạt
Motorbike riders, including ride-hailing drivers, waiting at an intersection in Vietnam

II. Problem and objectives

Background
Ride-hailing platforms record millions of GPS-based trips, but know little about who their riders are. Pick-up/drop-off locations, POIs and time of day carry strong hints about a rider's daily routine. This project uses 4.87M trips from 47,322 riders in Ho Chi Minh City and Hanoi (Jan–Aug 2026).
Motivation
Knowing a rider's segment enables better offer design, targeted promotions and service planning without asking riders for personal information. Behaviour-based tags are cheaper and more scalable than surveys, provided they are accurate and used responsibly.
Research question(s)
  1. Can ride behaviour alone classify a rider as Office WorkerParentor University Student?
  2. Which behavioural signals (locations, POIs, time windows) best separate the three segments?
  3. Which model (KNN, Decision Tree, Random Forest, Gradient Boosting) performs best, measured by macro-F1?

III. Data description

Source
Build Ride-Hailing raw data from xMap

IV. Data preprocessing

Preprocessing steps raw trips → rider-level features

StepActionJustification
Missing values
Train-median imputation; binary-mode imputation; invalid values → missing.Retains riders and handles unavailable measurements without learning from test data.
Outliers
Validity checks and train-fitted P1–P99 winsorization on six features.Limits extreme values while preserving rows, bounded shares and rare binary signals.
Encoding
Binary/share predictors; separate label encoding for user_tag.Converts categories to numeric inputs; keeps IDs and target labels outside X.
Scaling
Train-fitted z-scores for 60 features; two binary flags unchanged.Makes numeric scales comparable and applies consistent transformations to held-out data.
Feature engineering
62 rider-level POI, time, address, mobility, seasonal and spending features.Summarizes repeated travel behavior and trip context for rider classification.

Output · Lean data dictionary lean.csv · 1 row per rider · 64 columns

VariableTypeUnitDescription
rider_idID—Unique rider identifier (one row per rider)
user_tagLabel—Target class: Office Worker University Student Parent
dr_* / pu_* 16dr_universitypu_universitydr_cat_educationpu_cat_educationdr_education_centerdr_kindergartenpu_kindergartendr_under_high_schooldr_high_schooldr_beauty_academydr_office_buildingpu_office_buildingdr_cat_office_buildingpu_cat_office_buildingdr_any_poipu_any_poiNumericShare 0–1Share of trips whose drop / pickup POI is a given category or subtype (Education, University, Kindergarten, High School, Office Building…)
addr_* / dr_addr_* 15addr_dai_hocaddr_truong_ptaddr_officeaddr_chung_cuaddr_ngo_hemaddr_truong_ngoai_poiaddr_ktx_lien_quandr_addr_dai_hocdr_addr_truong_ptdr_addr_ktxdr_addr_officedr_addr_chung_cudr_addr_benh_viendr_addr_truong_ngoai_poidr_addr_office_ngoai_poiNumericShare 0–1Share of trips whose pickup / drop address text contains a keyword (đại học, THPT, tòa nhà, chung cư, ngõ, ký túc xá, bệnh viện…)
Home ↔ school / office flows 9nha_den_truongnha_den_officetu_truong_vedr_edu_ampu_edu_pmdr_edu_pmaddr_nha_den_truongaddr_nha_den_officeaddr_truong_ve_nhaNumericShare 0–1Trip direction patterns: home→school, school→home, home→office, plus school trips in morning (06–09h) and afternoon (15–18h)
sh_* time of day 6sh_am_peaksh_wd_vao_cash_wd_truash_wd_16hsh_wd_toish_nightNumericShare 0–1Share of trips booked in specific time windows (morning peak, lunch, 16h, evening, night), weekday vs all days
Anchor point 5neo_la_truong_ptneo_la_dai_hocsh_toi_diem_neosh_nha_diem_neotuan_co_diem_neoBinary / Numeric0/1, ShareRider’s most frequent destination cell: whether it is a school or university, and how often trips / weeks touch it
Service mix 3sh_carsh_foodsh_deli_wd_gio_hcNumericShare 0–1Share of trips by service: car, food, delivery during office hours
Seasonality 2tet_vs_hoc_kytruong_he_vs_hoc_kyNumericRatioTrip volume in Tết and summer break vs regular school months
Behavior & pricing 6active_dayspickup_diversitydriver_diversitydiscount_rateavg_surgeo_lai_median_gioNumericCount, Ratio, HoursActive days, pickup & driver diversity, discount rate, average surge, median dwell time between trips

V. Exploratory data analysis

All EDA uses the 80% Learning set only (37,858 riders), so the Test set stays unseen until final evaluation.
Office WorkerParentUniversity Student
Chart 1 · Class balance and split
Riders per user_tag, learning vs test
Insight: The classes are imbalanced (45 / 31 / 24%). Always guessing "Office Worker" already reaches 45% accuracy, so every split is stratified and models are selected by macro-F1 rather than accuracy.
Chart 2 · University drop-off signal
Distribution of dr_addr_dai_hoc by user_tag
Insight: 87.9% of students drop off at a university address on more than 14.9% of their trips, against 0.9% of office workers and 2.0% of parents. This one feature almost isolates the student class.
Chart 3 · Education drop-offs
Distribution of dr_cat_education by user_tag
Insight: 77.7% of parents end more than 26.4% of their trips at an education POI, against 18.6% of office workers. Students score high too, so this feature splits Parent from Office only after students are removed by Chart 2's signal.
Chart 4 · Time-of-day fingerprint
Share of trips by time window per user_tag
Insight: Each segment keeps its own schedule:
  • OfficeOffice workers ride more at night (16.0%) and on weekday evenings (12.8%), and order more deliveries in office hours.
  • ParentParents peak around the 16h school pick-up (8.2% vs 4.0% for office workers).
  • StudentStudents ride most at noon (15.0%).
Chart 5 · Segment fingerprints
Standardised mean difference heatmap
Insight: Each segment has a distinct anchor:
  • StudentStudents: +1.47 SD on university drop-offs.
  • ParentParents: +0.69 SD on a K-12 school anchor.
  • OfficeOffice workers: +0.46 SD on home → office trips, plus more car and night trips.
The features are behaviourally meaningful, not noise.
Chart 6 (optional) · Class separability
PCA projection coloured by user_tag
Insight: Students form their own cluster, while Office and Parent overlap in the middle. The first two components explain only 20.8% of variance, so non-linear models are needed. The overlap anticipates where most prediction errors fall.

VI. Methodology

We train four classifiers to guess each rider's user_tag from 60 behaviour features. To make the contest fair, all four use the same riders, the same features, the same practice rounds and the same final exam. Only the model and the settings it tries are different.

The workflow in five steps identical in all four scripts

1Split the riders

Within each tag, shuffle riders (seed 42). 80% go to Learning, 20% to Test.

Learning37,858
Test9,464
Learning only
2Clean features

Drop columns that never change or copy another column.

Features 62 → 60
Learning only
3Practice and tune

Try many settings. Score each one with 5-fold cross-validation.

Score = macro-F1
Learning only
4Retrain the winner

Fit the best setting again on all Learning riders.

One final model each
Test  9,464 riders stay locked. Nothing in steps 2–4 looks at them.
5Final exam, once

Predict the locked Test riders one time and compute the scores in section VII.

No retuning afterwards
Step 1 · Each tag is split on its own
Office Worker21,297 · test 4,259
17,038
Parent14,785 · test 2,957
11,828
University Student11,240 · test 2,248
8,992
Learning 80%Test 20%
So both sets keep the same mix of 45% Office, 31% Parent, 24% Student, and no rider appears in both.
Step 3 · 5-fold cross-validation = five practice exams
Train on ≈ 30,286Score on ≈ 7,572
Each setting is trained 5 times, each time scored on riders it did not see. Its CV score is the average of the 5. All models use the same 5 folds.

The four models from one simple model to teams of trees

Model 1 · Instance-based

K-Nearest Neighbours

CV macro-F10.884
? K = 11 nearest
The 11 most similar riders vote, and closer riders count more. Here: 7 Office, 3 Parent, 1 Student, so the answer is Office Worker.

Ask the neighbours. To tag a new rider, find the 11 riders in the Learning set whose behaviour is most similar and take their weighted vote. The model learns nothing in advance; it compares at prediction time, so features must first be put on the same scale.

SettingTriedChosen
Neighbours K30 values, 1 … 15111
Distancestraight-line or ManhattanManhattan
Voteequal, or closer counts morecloser counts more
ScalingStandard or Min-MaxStandard
48 settings triedpicked by 5-fold CV
Model 2 · Single tree

Decision Tree

CV macro-F10.871
University drop-off share high? yes no Student Education drop-off high? yes no Parent Office
Simplified view of the first two questions. The real tree asks up to 12 questions in a row and ends in 103 leaves.

A flowchart of yes/no questions. The tree picks the question that best separates the three tags, then repeats inside each branch. A tree that grows too far memorises the riders it was trained on, so it is cut back (pruned) until the cross-validation score is highest.

SettingTriedChosen
Split ruleGini or entropyGini
How to prunestop early (168 combos) or grow, then cut (26 levels)grow, then cut
Cut strength α0 to 0.010.000237
Final size12 levels, 103 leaves
194 settings triedpicked by 5-fold CV
Model 3 · Many trees in parallel

Random Forest

CV macro-F10.899
Learning 37,858 bootstrap ×500 tree 1tree 2… 500 average P(class) highest wins
Each tree learns from a different random resample of riders and sees only 19 of the 60 features at each question. The forest averages their answers.

Wisdom of the crowd. Grow 500 different trees, each from a random resample of riders and a random subset of features, then average their answers. Each tree makes different mistakes, so averaging cancels many of them.

SettingTriedChosen
Number of treesfixed500
Features per question7, 5 or 19 of 6019
Min riders per leaf1, 3 or 51
9 settings triedpicked by out-of-bag score
Model 4 · Many trees in sequence

Gradient Boosting Best

CV macro-F10.904
log-loss boosting rounds → best 30 rds stop validation (10%) train
Schematic. Training keeps improving, but the score on held-out riders stops improving. Training stops 30 rounds after the best point. Final model: 485 rounds.

Learn from mistakes, one small step at a time. Trees are added one after another. Each new small tree focuses on the riders the model still gets wrong and adds only a small correction. Training stops automatically when held-out riders stop improving.

SettingTriedChosen
Learning rate0.03, 0.1 or 0.30.03
Leaves per tree15, 31 or 6315
L2 penalty0 or 10
Roundsauto-stop, max 2,000485
18 settings triedpicked by 5-fold CV + auto-stop

How the models are scored on the 9,464 Test riders

MetricIn plain wordsUsed for
Macro-F1F1 for each tag, then the average. Small tags (Student, Parent) count as much as the 45% Office majority, which is why this is the main score.Choosing & ranking
95% CI of macro-F1Resample the Test riders 1,000 times. The range shows how much the score could move with a different Test sample.Reporting
AccuracyShare of riders tagged correctly. Always guessing “Office Worker” already gets 45%.Reporting
ROC-AUCHow well the predicted probabilities rank the true tag first. 0.5 = coin flip, 1 = perfect.Reporting
RMSE* and R²*How close the predicted probabilities are to the truth (1 for the true tag, 0 for the others). R² = 0 means no better than guessing the class shares.Supporting
*RMSE and R² are regression metrics, so here they are applied to probabilities.

VII. Results and evaluation

Best modelGradient Boostingbest on every metric
Test macro-F10.90495% CI 0.898–0.910 · baseline 0.207
Accuracy90.0%baseline 45.0%
ROC-AUC0.9780.5 = coin flip

Model comparison Test set, 9,464 riders · ranked by Test macro-F1

ModelAccuracyF1 (macro)F1 95% CICV F1RMSE*R²*ROC-AUCRank
Gradient Boosting0.9000.9040.898–0.9100.9040.2210.7720.9781
Random Forest0.8910.8950.888–0.9000.8990.2330.7470.9742
KNN0.8820.8840.877–0.8910.8840.2430.7250.9633
Decision Tree0.8620.8670.861–0.8740.8710.2640.6760.9554
Baseline (always “Office Worker”)0.4500.207––0.4630.0000.500–
*Classification task: RMSE and R² are computed on predicted probabilities vs one-hot labels (see VI). Lower RMSE and higher values elsewhere are better.
Performance chart · Test macro-F1 with 95% CI
0.860.870.880.890.900.91Gradient Boosting0.904CV 0.904Random Forest0.895CV 0.899KNN0.884CV 0.884Decision Tree0.867CV 0.871Macro-F1 (axis starts at 0.855; baseline = 0.207 is off the chart)
Test score with 95% bootstrap CI5-fold CV score
  • Teams of trees win. Random Forest beats a single tree by +0.027 F1. Boosting adds +0.009 over the forest; the paired bootstrap 95% CI of that gap is [+0.006, +0.013], so it is small but real.
  • No overfitting to the tuning. Every hollow CV dot sits less than 0.005 from its Test dot, and the order of the models is the same on CV and Test.
  • Where errors remain (Boosting F1 by tag): Student 0.953Office 0.906Parent 0.852 Office ↔ Parent mix-ups are 78% of all errors.
Validation strategy
  1. Hold-out: within each tag, 80% Learning / 20% Test (seed 42), saved to split/train.csv and split/test.csv and reused by all four scripts.
  2. Tune on Learning only: stratified 5-fold CV scored by macro-F1, with the same folds for every model. Random Forest uses its out-of-bag score; Boosting stops early on 10% of each training fold.
  3. Retrain the best setting on all Learning riders.
  4. Test once: accuracy, macro-F1 with a 1,000-resample bootstrap CI, ROC-AUC and probability RMSE / R².
Leakage controls
  • Test riders are never used for tuning.
  • KNN's scaler is refit inside every CV fold.
  • Duplicate columns are found on Learning only.
  • Early-stopping data comes from the training folds.
Evidence it worked
  • CV and Test macro-F1 differ by less than 0.005 for every model.
  • Same ranking on CV and Test, so the winner did not depend on Test.
  • The frozen split files make every rerun use the same riders.

VIII. Conclusions and limitations

RQ1
Can ride behaviour alone classify a rider?

Yes. The best model scores macro-F1 0.904 and 90.0% accuracy on 9,464 unseen riders, against 0.207 and 45.0% for always guessing Office Worker.

RQ2
Which signals separate the segments?

Where riders go, and when.

  • Studentuniversity drop-offs
  • Parentschool drop-offs, 16h pick-up
  • Officehome→office, night and car trips
RQ3
Which model performs best?

Gradient Boosting, first on every metric on both CV and Test. Random Forest is a close second (−0.009 F1).

Key takeaways
  • Students are easy; Office vs Parent is the real challenge. Student F1 is 0.953 but Parent only 0.852, and Office↔Parent mix-ups make up 78% of errors. Both groups commute on weekdays; a short school stop is often the only difference.
  • Combining trees pays off. Macro-F1 rises from 0.867 (one tree) to 0.895 (forest) to 0.904 (boosting).
  • The scores are stable. CV and Test differ by less than 0.005 and rank the models the same way, so results should hold for new riders from the same period.
  • A readable model is not far behind. A single tree reaches 0.867 with 103 if-then rules, a fair trade when decisions must be explained.
Limitations
  • Labels and features share a source. user_tag was assigned by matching stops to universities, schools and offices (section III), and the strongest features are drop-offs at the same POI types. Part of the score shows how well the model re-learns that rule, not true occupation or parenthood.
  • One tag per rider. A person can be both an office worker and a parent, so some Office/Parent errors cannot be avoided.
  • One time window, random split. The models were not tested on later months, so seasonal or behavioural drift is unmeasured.
  • Imperfect location features. POI matching uses a 25 m radius on noisy GPS, and three address features came out identical (only one was kept), so address text adds less signal than designed.
Potential biases
  • Coverage bias: POI and address quality vary by area, so riders in poorly mapped areas get weaker signals and likely more errors.
  • Selection bias: only riders who book trips are observed. Parents who do school runs on their own motorbike are missed, so “Parent” really means “parent who books school trips”.
  • Majority-class pull: Office Worker is 45% of riders, so ambiguous riders tend to be labelled Office, which hurts Parent most.
  • Proxy bias: “Parent” is inferred from doing the school run, which may reflect who in the household does it (often linked to gender) rather than parenthood itself.
Ethical considerations
  • Location trails are personal data. They reveal home, workplace and daily routine. Vietnam's Law on Personal Data Protection (No. 91/2025/QH15, in force since 1 Jan 2026, guided by Decree 356/2025, which replaced Decree 13/2023) requires a lawful basis or consent, purpose limitation and data minimisation.
  • Inferring life status is profiling. Use predicted tags for aggregate segmentation and offer design only. Never show them to riders or third parties, or use them for high-stakes individual decisions.
  • Children's routines: school drop-off patterns reveal where and when children are. Keep these features aggregated and access-restricted.
  • Fair treatment: if tags drive discounts, misclassified riders are treated differently. Monitor errors by segment, allow opt-out and document the model. Only pseudonymous rider_id was used.
Next steps
  1. Check labels against an independent source (e.g. a short rider survey).
  2. Test on later months (out-of-time validation).
  3. Target Office vs Parent with class weights, threshold tuning or multi-label tags.
  4. Try LightGBM / XGBoost and add SHAP explanations.