Day 6 [Python ML] 随机森林(Random Forests)演算法

前言

决策树(DecisionTree)如果leaf太多的话容易overfitting

若leaf太少的话则容易underfitting

但是随机森林(Random Forests)演算法 可以解决这个问题

并且随机森林(Random Forests)预设的参数就可以表现得很好了

范例

import pandas as pd
    
# 读取资料
melbourne_file_path = './Dataset/melb_data.csv'
melbourne_data = pd.read_csv(melbourne_file_path) 
# 处理缺失值
melbourne_data = melbourne_data.dropna(axis=0)
# 选择目标以及特徵
y = melbourne_data.Price
melbourne_features = ['Rooms', 'Bathroom', 'Landsize', 'BuildingArea', 
                        'YearBuilt', 'Lattitude', 'Longtitude']
X = melbourne_data[melbourne_features]

from sklearn.model_selection import train_test_split

# 切分资料成训练资料和验证资料，feature和target都要切分
# 这是使用随机切分的，设定random_stated可以确保每次切分的资料都是一样的 
train_X, val_X, train_y, val_y = train_test_split(X, y,random_state = 0)

在scikit-learn中，建立随机森林演算法的方法和决策树的方法是很像的

只是我们用的是RandomForestRegressor而不是DecisionTreeRegressor

from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_absolute_error

forest_model = RandomForestRegressor(random_state=1)
forest_model.fit(train_X, train_y)
melb_preds = forest_model.predict(val_X)
print(mean_absolute_error(val_y, melb_preds))