ID người dùng (ẩn danh), phân khúc, thành phố, ngày, giờ đi/đến, toạ độ đi/đến, POI/địa chỉ điểm đón và trả, quãng đường, thời gian di chuyển, phương tiện ước tính.
Chỉ cần đúng ở mức xu hướng và phân phối, không cần khớp từng chuyến. Ưu tiên độ phủ, quy mô panel đủ lớn và sự ổn định qua tám tháng.
| Variable | Type | Unit | Description |
|---|---|---|---|
| rider_id | ID | — | Unique rider identifier (one row per rider) |
| user_tag | Label | — | Target class: Office Worker University Student Parent |
dr_* / pu_* 16dr_universitypu_universitydr_cat_educationpu_cat_educationdr_education_centerdr_kindergartenpu_kindergartendr_under_high_schooldr_high_schooldr_beauty_academydr_office_buildingpu_office_buildingdr_cat_office_buildingpu_cat_office_buildingdr_any_poipu_any_poi | Numeric | Share 0–1 | Share of trips whose drop / pickup POI is a given category or subtype (Education, University, Kindergarten, High School, Office Building…) |
addr_* / dr_addr_* 15addr_dai_hocaddr_truong_ptaddr_officeaddr_chung_cuaddr_ngo_hemaddr_truong_ngoai_poiaddr_ktx_lien_quandr_addr_dai_hocdr_addr_truong_ptdr_addr_ktxdr_addr_officedr_addr_chung_cudr_addr_benh_viendr_addr_truong_ngoai_poidr_addr_office_ngoai_poi | Numeric | Share 0–1 | Share of trips whose pickup / drop address text contains a keyword (đại học, THPT, tòa nhà, chung cư, ngõ, ký túc xá, bệnh viện…) |
Home ↔ school / office flows 9nha_den_truongnha_den_officetu_truong_vedr_edu_ampu_edu_pmdr_edu_pmaddr_nha_den_truongaddr_nha_den_officeaddr_truong_ve_nha | Numeric | Share 0–1 | Trip direction patterns: home→school, school→home, home→office, plus school trips in morning (06–09h) and afternoon (15–18h) |
sh_* time of day 6sh_am_peaksh_wd_vao_cash_wd_truash_wd_16hsh_wd_toish_night | Numeric | Share 0–1 | Share of trips booked in specific time windows (morning peak, lunch, 16h, evening, night), weekday vs all days |
Anchor point 5neo_la_truong_ptneo_la_dai_hocsh_toi_diem_neosh_nha_diem_neotuan_co_diem_neo | Binary / Numeric | 0/1, Share | Rider’s most frequent destination cell: whether it is a school or university, and how often trips / weeks touch it |
Service mix 3sh_carsh_foodsh_deli_wd_gio_hc | Numeric | Share 0–1 | Share of trips by service: car, food, delivery during office hours |
Seasonality 2tet_vs_hoc_kytruong_he_vs_hoc_ky | Numeric | Ratio | Trip volume in Tết and summer break vs regular school months |
Behavior & pricing 6active_dayspickup_diversitydriver_diversitydiscount_rateavg_surgeo_lai_median_gio | Numeric | Count, Ratio, Hours | Active days, pickup & driver diversity, discount rate, average surge, median dwell time between trips |
| Step | Action | Justification |
|---|---|---|
Missing values | Train-median imputation; binary-mode imputation; invalid values → missing. | Retains riders and handles unavailable measurements without learning from test data. |
| ||
Outliers | Validity checks and train-fitted P1–P99 winsorization on six features. | Limits extreme values while preserving rows, bounded shares and rare binary signals. |
Winsorization áp dụng cho đúng sáu feature: active_daysavg_surgediscount_ratetet_vs_hoc_kytruong_he_vs_hoc_kyo_lai_median_gio
| ||
Encoding | Binary/share predictors; separate label encoding for user_tag. | Converts categories to numeric inputs; keeps IDs and target labels outside X. |
| ||
Scaling | Train-fitted z-scores for 60 features; two binary flags unchanged. | Makes numeric scales comparable and applies consistent transformations to held-out data. |
z = (x − mean_train) / std_train
| ||
Feature engineering | 62 rider-level POI, time, address, mobility, seasonal and spending features. | Summarizes repeated travel behavior and trip context for rider classification. |
| ||






| Model / Technique | Purpose | Key hyperparameters (searched → selected) | Notes |
|---|---|---|---|
| Models | |||
| Baseline (majority class) | Lower bound that any useful model must beat | None (always predicts Office Worker) | Test accuracy 0.450, macro-F1 0.207 |
| K-Nearest Neighbours | Distance-based classifier: a rider takes the majority tag of the K most similar riders | K ∈ {1, 3, …, 51, 61, 75, 101, 151} → 11; distance Euclidean/Manhattan → Manhattan; weights → distance; scaler Standard/MinMax → Standard | Scaler sits inside the pipeline, so there is no leakage. Manhattan wins because several features are heavy-tailed. Slowest to predict (4.6 s for 9,464 riders). |
| Decision Tree (CART) | Transparent if-then rules; shows which splits matter | criterion × max_depth × min_samples_leaf (168 combos), then post-pruning ccp_alpha (26 values) → Gini, α = 2.37e-4 (depth 12, 103 leaves) | No scaling needed. Root split dr_addr_dai_hoc ≤ 0.149 gives an information gain of 0.555 bits. |
| Random Forest | Bagging with random feature subsets to cut single-tree variance | n_estimators = 500; max_features √m / log₂m / m/3 → m/3 (19); min_samples_leaf {1, 3, 5} → 1 | Tuned on out-of-bag (OOB) score: OOB F1 0.899 matches CV F1 0.899. OOB error is flat after about 300 trees. |
| Gradient Boosting (HistGradientBoosting)Best | Sequential shallow trees, each fitting the gradient of the log-loss; usually the strongest model on tabular data | learning_rate {0.03, 0.1, 0.3} → 0.03; max_leaf_nodes {15, 31, 63} → 15; L2 {0, 1} → 0; rounds chosen by early stopping → 485 | With lr = 0.3 the validation loss bottoms out after 14 rounds and then diverges, so early stopping is essential. |
| Techniques | |||
| Stratified rider split | Honest hold-out estimate | Within each user_tag: shuffle riders once (seed 42), 80% Learning / 20% Test | 37,858 / 9,464 riders, no overlap. Frozen in split/train.csv and split/test.csv and reused on every run. |
| Stratified 5-fold CV + grid search | Choose hyperparameters without touching Test | Scoring = macro-F1; the same folds for every model | Test set is used once, at the end |
| Feature cleaning | Remove redundancy and avoid double weighting in distances | Drop constant or exactly duplicated columns, identified on the Learning set | 62 → 60 features: 2 columns were exact copies of addr_ngo_hem |
| Bootstrap CI and permutation importance | Quantify uncertainty; explain drivers | 1,000 bootstrap resamples; 3 permutation repeats | Used for interpretation only, never for model selection |
| Model | Accuracy | F1 (macro) | RMSE* | R²* | ROC-AUC | Rank |
|---|---|---|---|---|---|---|
| Gradient Boosting | 0.900 | 0.904 | 0.221 | 0.772 | 0.978 | 1 |
| Random Forest | 0.891 | 0.895 | 0.233 | 0.747 | 0.974 | 2 |
| KNN | 0.882 | 0.884 | 0.243 | 0.725 | 0.963 | 3 |
| Decision Tree | 0.862 | 0.867 | 0.264 | 0.676 | 0.955 | 4 |
| Baseline (majority class) | 0.450 | 0.207 | 0.463 | 0.000 | 0.500 | – |

split/train.csv and split/test.csv, and every rerun and all four model scripts reuse these exact files.