Data Science Group Project

I. Title slide

Predicting Rider Persona Segments from Ride-Hailing Behaviour

Members

01Nguyễn Giao Chi
02Trần Anh Minh
03Huỳnh Ngọc Tấn Thuận
04Huỳnh Trọng Nghĩa
05Nguyễn Quốc Khánh
06Ngô Quốc Đạt

II. Problem and objectives

Background
Ride-hailing platforms record millions of GPS-based trips, but know little about who their riders are. Pick-up/drop-off locations, POIs and time of day carry strong hints about a rider's daily routine. This project uses 4.87M trips from 47,322 riders in Ho Chi Minh City and Hanoi (Jan–Aug 2026).
Motivation
Knowing a rider's segment enables better offer design, targeted promotions and service planning without asking riders for personal information. Behaviour-based tags are cheaper and more scalable than surveys, provided they are accurate and used responsibly.
Research question(s)
  1. Can ride behaviour alone classify a rider as Office WorkerParentor University Student?
  2. Which behavioural signals (locations, POIs, time windows) best separate the three segments?
  3. Which model (KNN, Decision Tree, Random Forest, Gradient Boosting) performs best, measured by macro-F1?

III. Data description

Source
Build Ride-Hailing raw data from xMap

Data dictionary

VariableTypeUnitDescription
rider_idID—Unique rider identifier (one row per rider)
user_tagLabel—Target class: Office Worker University Student Parent
dr_* / pu_* 16dr_universitypu_universitydr_cat_educationpu_cat_educationdr_education_centerdr_kindergartenpu_kindergartendr_under_high_schooldr_high_schooldr_beauty_academydr_office_buildingpu_office_buildingdr_cat_office_buildingpu_cat_office_buildingdr_any_poipu_any_poiNumericShare 0–1Share of trips whose drop / pickup POI is a given category or subtype (Education, University, Kindergarten, High School, Office Building…)
addr_* / dr_addr_* 15addr_dai_hocaddr_truong_ptaddr_officeaddr_chung_cuaddr_ngo_hemaddr_truong_ngoai_poiaddr_ktx_lien_quandr_addr_dai_hocdr_addr_truong_ptdr_addr_ktxdr_addr_officedr_addr_chung_cudr_addr_benh_viendr_addr_truong_ngoai_poidr_addr_office_ngoai_poiNumericShare 0–1Share of trips whose pickup / drop address text contains a keyword (đại học, THPT, tòa nhà, chung cư, ngõ, ký túc xá, bệnh viện…)
Home ↔ school / office flows 9nha_den_truongnha_den_officetu_truong_vedr_edu_ampu_edu_pmdr_edu_pmaddr_nha_den_truongaddr_nha_den_officeaddr_truong_ve_nhaNumericShare 0–1Trip direction patterns: home→school, school→home, home→office, plus school trips in morning (06–09h) and afternoon (15–18h)
sh_* time of day 6sh_am_peaksh_wd_vao_cash_wd_truash_wd_16hsh_wd_toish_nightNumericShare 0–1Share of trips booked in specific time windows (morning peak, lunch, 16h, evening, night), weekday vs all days
Anchor point 5neo_la_truong_ptneo_la_dai_hocsh_toi_diem_neosh_nha_diem_neotuan_co_diem_neoBinary / Numeric0/1, ShareRider’s most frequent destination cell: whether it is a school or university, and how often trips / weeks touch it
Service mix 3sh_carsh_foodsh_deli_wd_gio_hcNumericShare 0–1Share of trips by service: car, food, delivery during office hours
Seasonality 2tet_vs_hoc_kytruong_he_vs_hoc_kyNumericRatioTrip volume in Tết and summer break vs regular school months
Behavior & pricing 6active_dayspickup_diversitydriver_diversitydiscount_rateavg_surgeo_lai_median_gioNumericCount, Ratio, HoursActive days, pickup & driver diversity, discount rate, average surge, median dwell time between trips

IV. Data preprocessing

StepActionJustification
Missing values
Train-median imputation; binary-mode imputation; invalid values → missing.Retains riders and handles unavailable measurements without learning from test data.
Outliers
Validity checks and train-fitted P1–P99 winsorization on six features.Limits extreme values while preserving rows, bounded shares and rare binary signals.
Encoding
Binary/share predictors; separate label encoding for user_tag.Converts categories to numeric inputs; keeps IDs and target labels outside X.
Scaling
Train-fitted z-scores for 60 features; two binary flags unchanged.Makes numeric scales comparable and applies consistent transformations to held-out data.
Feature engineering
62 rider-level POI, time, address, mobility, seasonal and spending features.Summarizes repeated travel behavior and trip context for rider classification.

V. Exploratory data analysis

All EDA uses the 80% Learning set only (37,858 riders), so the Test set stays unseen until final evaluation.
Office WorkerParentUniversity Student
Chart 1 · Class balance and split
Riders per user_tag, learning vs test
Insight: The classes are imbalanced (45 / 31 / 24%). Always guessing "Office Worker" already reaches 45% accuracy, so every split is stratified and models are selected by macro-F1 rather than accuracy.
Chart 2 · University drop-off signal
Distribution of dr_addr_dai_hoc by user_tag
Insight: 87.9% of students drop off at a university address on more than 14.9% of their trips, against 0.9% of office workers and 2.0% of parents. This one feature almost isolates the student class.
Chart 3 · Education drop-offs
Distribution of dr_cat_education by user_tag
Insight: 77.7% of parents end more than 26.4% of their trips at an education POI, against 18.6% of office workers. Students score high too, so this feature splits Parent from Office only after students are removed by Chart 2's signal.
Chart 4 · Time-of-day fingerprint
Share of trips by time window per user_tag
Insight: Each segment keeps its own schedule:
  • OfficeOffice workers ride more at night (16.0%) and on weekday evenings (12.8%), and order more deliveries in office hours.
  • ParentParents peak around the 16h school pick-up (8.2% vs 4.0% for office workers).
  • StudentStudents ride most at noon (15.0%).
Chart 5 · Segment fingerprints
Standardised mean difference heatmap
Insight: Each segment has a distinct anchor:
  • StudentStudents: +1.47 SD on university drop-offs.
  • ParentParents: +0.69 SD on a K-12 school anchor.
  • OfficeOffice workers: +0.46 SD on home → office trips, plus more car and night trips.
The features are behaviourally meaningful, not noise.
Chart 6 (optional) · Class separability
PCA projection coloured by user_tag
Insight: Students form their own cluster, while Office and Parent overlap in the middle. The first two components explain only 20.8% of variance, so non-linear models are needed. The overlap anticipates where most prediction errors fall.

VI. Methodology

Model / TechniquePurposeKey hyperparameters (searched → selected)Notes
Models
Baseline (majority class)Lower bound that any useful model must beatNone (always predicts Office Worker)Test accuracy 0.450, macro-F1 0.207
K-Nearest NeighboursDistance-based classifier: a rider takes the majority tag of the K most similar ridersK ∈ {1, 3, …, 51, 61, 75, 101, 151} → 11; distance Euclidean/Manhattan → Manhattan; weights → distance; scaler Standard/MinMax → StandardScaler sits inside the pipeline, so there is no leakage. Manhattan wins because several features are heavy-tailed. Slowest to predict (4.6 s for 9,464 riders).
Decision Tree (CART)Transparent if-then rules; shows which splits mattercriterion × max_depth × min_samples_leaf (168 combos), then post-pruning ccp_alpha (26 values) → Gini, α = 2.37e-4 (depth 12, 103 leaves)No scaling needed. Root split dr_addr_dai_hoc ≤ 0.149 gives an information gain of 0.555 bits.
Random ForestBagging with random feature subsets to cut single-tree variancen_estimators = 500; max_features √m / log₂m / m/3 → m/3 (19); min_samples_leaf {1, 3, 5} → 1Tuned on out-of-bag (OOB) score: OOB F1 0.899 matches CV F1 0.899. OOB error is flat after about 300 trees.
Gradient Boosting (HistGradientBoosting)BestSequential shallow trees, each fitting the gradient of the log-loss; usually the strongest model on tabular datalearning_rate {0.03, 0.1, 0.3} → 0.03; max_leaf_nodes {15, 31, 63} → 15; L2 {0, 1} → 0; rounds chosen by early stopping → 485With lr = 0.3 the validation loss bottoms out after 14 rounds and then diverges, so early stopping is essential.
Techniques
Stratified rider splitHonest hold-out estimateWithin each user_tag: shuffle riders once (seed 42), 80% Learning / 20% Test37,858 / 9,464 riders, no overlap. Frozen in split/train.csv and split/test.csv and reused on every run.
Stratified 5-fold CV + grid searchChoose hyperparameters without touching TestScoring = macro-F1; the same folds for every modelTest set is used once, at the end
Feature cleaningRemove redundancy and avoid double weighting in distancesDrop constant or exactly duplicated columns, identified on the Learning set62 → 60 features: 2 columns were exact copies of addr_ngo_hem
Bootstrap CI and permutation importanceQuantify uncertainty; explain drivers1,000 bootstrap resamples; 3 permutation repeatsUsed for interpretation only, never for model selection

VII. Results and evaluation

Model comparison Test set, 9,464 riders

ModelAccuracyF1 (macro)RMSE*R²*ROC-AUCRank
Gradient Boosting0.9000.9040.2210.7720.9781
Random Forest0.8910.8950.2330.7470.9742
KNN0.8820.8840.2430.7250.9633
Decision Tree0.8620.8670.2640.6760.9554
Baseline (majority class)0.4500.2070.4630.0000.500–
*This is a classification task, so RMSE and R² are computed on the predicted class probabilities against one-hot labels.
Performance chart
Test macro-F1 with bootstrap CI and CV score per model
  • Boosting vs Random Forest: +0.009 F1, paired bootstrap 95% CI [+0.006, +0.013]. The gap is small but excludes 0.
  • Ensembles beat single models: the forest gains +0.027 F1 over one tree.
  • By class (Boosting F1): Student 0.953Office 0.906Parent 0.852 Office↔Parent confusions make up 78% of all errors.
Validation strategy
  1. Hold-out: within each user_tag, shuffle riders once (seed 42) and keep 80% for Learning and 20% for Test. The split is saved as split/train.csv and split/test.csv, and every rerun and all four model scripts reuse these exact files.
  2. Tuning: stratified 5-fold CV on the Learning set, with grid search scored by macro-F1. The random forest also uses its OOB score, and boosting picks its number of rounds by early stopping on 10% of each training fold.
  3. Refit the best configuration on the full Learning set.
  4. Test once: report accuracy, macro-F1, probability-based RMSE and R², and ROC-AUC, with 95% bootstrap CIs.
Leakage controls
  • Scaling is fitted inside each CV fold (pipeline).
  • Duplicate columns are detected on the Learning set only.
  • Early-stopping data is carved from the training folds.
Evidence it worked
  • CV and Test macro-F1 differ by at most 0.004 for every model.
  • Rerunning the pipeline and the four standalone model scripts reproduced every metric exactly, on the same frozen split.

VIII. Conclusions and limitations

Key takeaways
  • Ride behaviour predicts user_tag well. Gradient Boosting reaches macro-F1 0.904, 90.0% accuracy and AUC 0.978, against 0.207 F1 for the baseline.
  • Location semantics carry the signal. University drop-offs identify students, and school drop-offs and the 16h pick-up separate parents from office workers.
  • Ensembles win, and the tuning generalises. Boosting > Random Forest > KNN > Tree holds on CV and on Test, with a gap of at most 0.004.
  • Remaining error is concentrated. Office↔Parent confusions are 78% of mistakes, so improving that pair is the next target.
Limitations
  • Single-label assumption: a rider can be both an office worker and a parent, so the Office/Parent overlap is partly real.
  • No out-of-time test: one data window and a random split. Performance on future months or other seasons (Tết, summer break) is untested.
  • Feature quality depends on POI matching (25 m radius, GPS noise). Two columns were exact duplicates, which suggests an upstream pipeline issue.
  • Search space was finite: no class weights, threshold tuning or XGBoost/LightGBM yet. RMSE and R² are probability-based proxies, not regression metrics.
Potential biases
  • Label bias: if user_tag was assigned by rules built on the same POI signals, scores are inflated because the model re-learns the rule. The labelling method should be confirmed.
  • Coverage bias: POI coverage and address quality differ across areas, so riders in poorly mapped areas get weaker signals and likely more errors.
  • Selection bias: only riders active on Be are included. Light users, and parents who do school runs on their own vehicle, are under-represented.
  • Class-prior and proxy bias:
    • The 45% Office majority pulls ambiguous riders toward Office (KNN Parent recall is only 0.781).
    • "Parent" is effectively inferred from doing the school run, which may track household roles such as gender.
Ethical considerations
  • Location data is sensitive personal data under Vietnam's Decree 13/2023/ND-CP. Use requires a lawful basis or consent, purpose limitation and data minimisation.
  • Inferring life status is profiling. Predicted tags should drive aggregate segmentation and offer design, never be exposed to riders or third parties, and never drive high-stakes individual decisions.
  • Children's routines: kindergarten and school drop-off patterns reveal where children are and when. Keep them aggregated and access-restricted.
  • Fair use: if tags drive discounts, a misclassified rider gets different treatment. Monitor error rates by segment, allow opt-out, and document the model (model card). Only pseudonymous rider_id was used.

IX. References

Data sources
Cited materials