The aim is to classify individuals into different obesity risk categories based on multiple input features such as physical activity, eating habits, age, gender, and other health indicators. This challenge was hosted on Kaggle as part of Playground Series - Season 4, Episode 2. Submissions were evaluated based on the accuracy score .
You can explore the complete methodology in this notebook: 🔗 PS4E2 - EDA LGBM XGB CAT Blend
Key steps followed:
-
📊 Exploratory Data Analysis (EDA):
- Assessed feature distributions and relationships.
- Checked for outliers, imbalance, and preprocessing needs.
-
🧠 Model Training:
- Trained three models independently: LightGBM, XGBoost, and CatBoost.
- Tuned hyperparameters for optimized performance.
-
🔀 Model Blending:
- Combined predictions using a weighted average strategy to boost performance.
- Aimed to reduce individual model bias and variance.
-
✅ Public Leaderboard:
- Achieved 91.40% and 91.54% accuracy scores.
-
🏁 Private Leaderboard:
- Best score of 90.89% on final submission.
-
🥇 Rank Achieved:
- Ranked 364 / 3746 participants and 3587 teams, as a solo participant.
- 📂 Kaggle Competition: Multi-Class Prediction of Obesity Risk
- 📁 Dataset: Data Info
- 📊 Data Source: Obesity or CVD risk
- Language: Python 🐍
- Libraries:
pandas,numpyfor data handlingmatplotlib,seabornfor visualizationlightgbm,xgboost,catboostfor modeling
- Tools:
- Jupyter Notebook 📓 for development and analysis
- Colab/kaggle kernels
📌 This project demonstrates the impact of ensemble modeling and feature understanding in achieving high performance on multi-class classification tasks in the health domain.