Machine Learning · Intelligent Transportation Systems

Vehicle Type Classification Using NGSIM US-101

Developed and evaluated supervised machine-learning models for classifying motorcycles, passenger cars, and trucks using the NGSIM US-101 traffic trajectory dataset. Models compared: Logistic Regression, Decision Tree, and Random Forest.

System overview

Functional overview
New explanatory diagram: functional areas grouped by responsibility. This is not an as-built schematic.

Author

Mohammed Mahyoub

Method

Vehicle records were aggregated to the vehicle level to reduce data leakage and create independent observations for training and evaluation.

Scope

Data preprocessing, exploratory analysis, feature engineering, model comparison, hyperparameter optimisation, classification evaluation, and feature-importance analysis.

Engineering rationale

Why vehicle-level aggregation matters

A vehicle appears in many trajectory frames. Randomly splitting those frames can put the same vehicle into both training and testing. The notebook aggregates records into vehicle observations before modelling so the evaluation unit matches the classification task.

Features and model choices

Vehicle length and width capture physical dimensions; mean speed, mean acceleration, dominant lane and mean space headway describe traffic behaviour. Logistic Regression, Decision Tree and Random Forest provide contrasting decision boundaries and interpretability.

Evaluation beyond headline accuracy

Motorcycles, passenger cars and trucks form an imbalanced classification problem. Macro-F1 and classwise recall are needed alongside accuracy. The dimensions-only versus full-feature comparison asks how much behaviour adds beyond vehicle size. Numerical results belong to the notebook outputs and should retain their experimental context.

Reproducibility

The repository contains the Jupyter notebook and requirements.txt. Run it with the original dataset paths reviewed first; preserve vehicle identity boundaries during train/test preparation and tuning. The Zenodo DOI identifies the archived research output, while GitHub contains the current documentation.

Engineering workflow
Explanatory workflow; read left to right across each row.

Evidence and validation

Suggested evidence checks for reviewing this work. These are not claimed pass results.

Full documentation and sources