上采样与下采样_sklearn 采样

作者：空白诗007 | 2024-08-20 07:06:50

踩

sklearn 采样

数据分析中的上采样和下采样

背景： 在分类问题中，由于各种原因，我们所获取到的数据集很容易出现正负样本的不平衡，或者某些数据特别多，有些数据则特别少，在这样的数据集中，进行训练，我们很难获得较好的训练结果，为此我们引入了上采样与下采样解决样本不平衡的问题。

解决方法，采用Imblearn包

pip install imblearn
1

上采样（把数量少的增加到与数量多的相近）

随机下采样

#使用make_classification生成样本数据
from sklearn.datasets import make_classification
X, y = make_classification(n_samples=5000, n_features=2, n_informative=2,
                            n_redundant=0, n_repeated=0, n_classes=3,
                            n_clusters_per_class=1,
                           weights=[0.01, 0.05, 0.94],
                            class_sep=0.8, random_state=0)
1
2
3
4
5
6
7

SMOTE模型

对于少数类中每一个样本x，以欧氏距离为标准计算它到少数类样本集中所有样本的距离，得到其k近邻
根据样本不平衡比例设置一个确定采样倍率N，对于每一个少数类样本x，从其k近邻中随机选择若干个样本，假设选择的近邻为xn
对于每一个随机选出来的近邻xn，分别与原样本按照如下的公式构建新的样本

#SMOTE过采样
from imblearn.over_sampling import SMOTE, ADASYN
X_resampled, y_resampled = SMOTE().fit_resample(X, y)
print(sorted(Counter(y_resampled).items()))
1
2
3
4

ADASYN

#ADASYN过采样
X_resampled, y_resampled = ADASYN().fit_resample(X, y)
print(sorted(Counter(y_resampled).items()))
1
2
3

SMOTE: 对于少数类样本a, 随机选择一个最近邻的样本b, 然后从a与b的连线上随机选取一个点c作为新的少数类样本;
ADASYN: 关注的是在那些基于K近邻分类器被错误分类的原始样本附近生成新的少数类样本；

下采样（把数量多的减少到与数量少的相近）

在这里插入图片描述

Prototype generation

#欠采样
from imblearn.under_sampling import ClusterCentroids
cc = ClusterCentroids(random_state=0)
X_resampled, y_resampled = cc.fit_resample(X, y)
print(sorted(Counter(y_resampled).items()))
1
2
3
4
5

随机下采样

#随机欠采样
from imblearn.under_sampling import RandomUnderSampler
rus = RandomUnderSampler(random_state=0)
X_resampled, y_resampled = rus.fit_resample(X, y)
print(sorted(Counter(y_resampled).items()))
1
2
3
4
5

组合采样

#两种组合采样的方法
from collections import Counter
from sklearn.datasets import make_classification
X, y = make_classification(n_samples=5000, n_features=2, n_informative=2,
                           n_redundant=0, n_repeated=0, n_classes=3,
                           n_clusters_per_class=1,
                           weights=[0.01, 0.05, 0.94],
                           class_sep=0.8, random_state=0)
print(sorted(Counter(y).items()))
 
from imblearn.combine import SMOTEENN
smote_enn = SMOTEENN(random_state=0)
X_resampled, y_resampled = smote_enn.fit_resample(X, y)
print(sorted(Counter(y_resampled).items()))
 
from imblearn.combine import SMOTETomek
smote_tomek = SMOTETomek(random_state=0)
X_resampled, y_resampled = smote_tomek.fit_resample(X, y)
print(sorted(Counter(y_resampled).items()))
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19

图像处理中的上采样和下采样

上采样

下采样

声明：本文内容由网友自发贡献，不代表【wpsshop博客】立场，版权归原作者所有，本站不承担相应法律责任。如您发现有侵权的内容，请联系我们。转载请注明出处：https://www.wpsshop.cn/w/空白诗007/article/detail/1005798