Simple Machine Learning for Weather Classification
In this project, I built a machine-learning model to classify weather conditions from features such as temperature, wind speed, atmospheric pressure, and UV index. The model uses XGBoost as its classification algorithm.
XGBoost, or Extreme Gradient Boosting, is an open-source machine-learning library based on decision trees. It supports classification, regression, and ranking, and can train on large datasets without licensing costs.
Benefits of XGBoost:
- Execution speed: XGBoost is optimized for efficient training and prediction, which is valuable when working with large datasets.
- Model performance: It often performs competitively against methods such as random forests, gradient boosting machines (GBM), and gradient-boosted decision trees (GBDT).
For more information, see the official XGBoost documentation.
The project follows these steps:
-
Import the supporting libraries.
pythonfrom sklearn.model_selection import train_test_split from sklearn.metrics import accuracy_score from sklearn.preprocessing import LabelEncoder import numpy as np import pandas as pd import xgboost as xgb import os -
Load the dataset used to train the model. The dataset is available in this GitHub repository.
pythondf = pd.read_csv('weather_classification_data.csv') -
Initialize an encoder to convert categorical values into numeric values.
pythonencoder = LabelEncoder() -
Transform the categorical features. In this dataset, the categorical columns are Cloud Cover, Location, and Season.
pythondf['Cloud Cover'] = encoder.fit_transform(df['Cloud Cover']) df['Location'] = encoder.fit_transform(df['Location']) df['Season'] = encoder.fit_transform(df['Season']) -
Transform the target column because the weather label is also stored as a string.
pythony_target = encoder.fit_transform(df["Weather Type"]) -
Separate the feature data from the target data.
pythondataset = df[[i for i in df.columns[:-1]]] -
Split the dataset into training and test sets. This example reserves 20% for testing with
test_size=0.2.pythonx_train, x_test, y_train, y_test = train_test_split( dataset, y_target, test_size=0.2, random_state=50 ) -
Define the model parameters.
objective: 'multi:softmax'selects multiclass classification using the softmax objective.num_class: len(df.columns[:-1])sets the number of classes.df.columns[:-1]selects every column except the final target column in this example.eta: 0.01sets the learning rate. A small value can make training slower but more stable.max_depth: 6controls the maximum depth of each decision tree. Larger values increase model capacity but can also increase the risk of overfitting.
pythonparameters = { 'objective':'multi:softmax', 'num_class': len(df.columns[:-1]), 'eta': 0.01, 'max_depth': 6, } -
Create the XGBoost classifier and save the trained model to
XGBClassifier.json.pythonmodel = xgb.XGBClassifier( parameters ) model.fit(x_train, y_train) model.save_model("XGBClassifier.json") -
Evaluate the model on the test set. An accuracy closer to 1 indicates stronger performance on this dataset, while a value below 0.5 may indicate that the model needs further improvement.
pythonpred = model.predict(x_test) accuracy = accuracy_score(y_test, pred) print("Accuracy:", accuracy)