arrow_backAll Posts·5 Min Read·2025-02-14

Simple Machine Learning for Weather Classification

Artificial Intelligence (AI)XGBoostClassification

In this project, I built a machine-learning model to classify weather conditions from features such as temperature, wind speed, atmospheric pressure, and UV index. The model uses XGBoost as its classification algorithm.

XGBoost, or Extreme Gradient Boosting, is an open-source machine-learning library based on decision trees. It supports classification, regression, and ranking, and can train on large datasets without licensing costs.

Benefits of XGBoost:

  • Execution speed: XGBoost is optimized for efficient training and prediction, which is valuable when working with large datasets.
  • Model performance: It often performs competitively against methods such as random forests, gradient boosting machines (GBM), and gradient-boosted decision trees (GBDT).

For more information, see the official XGBoost documentation.

The project follows these steps:

  • Import the supporting libraries.

    python
    from sklearn.model_selection import train_test_split
    from sklearn.metrics import accuracy_score
    from sklearn.preprocessing import LabelEncoder
    import numpy as np
    import pandas as pd
    import xgboost as xgb
    import os
    
  • Load the dataset used to train the model. The dataset is available in this GitHub repository.

    python
    df = pd.read_csv('weather_classification_data.csv')
    
  • Initialize an encoder to convert categorical values into numeric values.

    python
    encoder = LabelEncoder()
    
  • Transform the categorical features. In this dataset, the categorical columns are Cloud Cover, Location, and Season.

    python
    df['Cloud Cover'] = encoder.fit_transform(df['Cloud Cover'])
    df['Location'] = encoder.fit_transform(df['Location'])
    df['Season'] = encoder.fit_transform(df['Season'])
    
  • Transform the target column because the weather label is also stored as a string.

    python
    y_target = encoder.fit_transform(df["Weather Type"])
    
  • Separate the feature data from the target data.

    python
    dataset = df[[i for i in df.columns[:-1]]]
    
  • Split the dataset into training and test sets. This example reserves 20% for testing with test_size=0.2.

    python
    x_train, x_test, y_train, y_test = train_test_split(
      dataset, 
      y_target,
      test_size=0.2,
      random_state=50
    )
    
  • Define the model parameters.

    • objective: 'multi:softmax' selects multiclass classification using the softmax objective.
    • num_class: len(df.columns[:-1]) sets the number of classes. df.columns[:-1] selects every column except the final target column in this example.
    • eta: 0.01 sets the learning rate. A small value can make training slower but more stable.
    • max_depth: 6 controls the maximum depth of each decision tree. Larger values increase model capacity but can also increase the risk of overfitting.
    python
    parameters = {
        'objective':'multi:softmax',
        'num_class': len(df.columns[:-1]),
        'eta': 0.01,
        'max_depth': 6,
    }
    
  • Create the XGBoost classifier and save the trained model to XGBClassifier.json.

    python
    model = xgb.XGBClassifier(
        parameters
    )
    model.fit(x_train, y_train)
    model.save_model("XGBClassifier.json")
    
  • Evaluate the model on the test set. An accuracy closer to 1 indicates stronger performance on this dataset, while a value below 0.5 may indicate that the model needs further improvement.

    python
    pred = model.predict(x_test)
    accuracy = accuracy_score(y_test, pred)
    print("Accuracy:", accuracy)