Python产生batch数据的方法 - 代码天地

Python产生batch数据的方法

编程语言 2018-10-27 04:59:54 阅读次数: 0

版权声明：本文为博主原创文章，未经博主允许不得转载。 https://blog.csdn.net/huanghaocs/article/details/83242353

参考此文：https://blog.csdn.net/qq_33039859/article/details/79901667

产生batch数据

输入data中每个样本可以有多个特征，和一个标签，最好都是numpy.array格式。
datas = [data1, data2, …, dataN ], labels = [label1, label2, …, labelN]，
其中data[i] = [feature1, feature2,…featureM], 表示每个样本数据有M个特征。
输入我们方法的数据，all_data = [datas, labels] 。

代码实现

通过索引值来产生batch大小的数据，同时提供是否打乱顺序的选择，根据随机产生数据量范围类的索引值来打乱顺序。

import numpy as np

def batch_generator(all_data , batch_size, shuffle=True):
    """
    :param all_data : all_data整个数据集
    :param batch_size: batch_size表示每个batch的大小
    :param shuffle: 每次是否打乱顺序
    :return:
    """
    all_data = [np.array(d) for d in all_data]
    data_size = all_data[0].shape[0]
    print("data_size: ", data_size)
    if shuffle:
        p = np.random.permutation(data_size)
        all_data = [d[p] for d in all_data]

    batch_count = 0
    while True:
        if batch_count * batch_size + batch_size > data_size:
            batch_count = 0
            if shuffle:
                p = np.random.permutation(data_size)
                all_data = [d[p] for d in all_data]
        start = batch_count * batch_size
        end = start + batch_size
        batch_count += 1
        yield [d[start: end] for d in all_data]

测试数据

样本数据x和标签y可以分开输入，也可以同时输入。

# 输入x表示有23个样本，每个样本有两个特征
# 输出y表示有23个标签，每个标签取值为0或1
x = np.random.random(size=[23, 2])
y = np.random.randint(2, size=[23,1])

batch_size = 5
batch_gen = batch_generator([x, y],  batch_size)
for i in range(20):
    batch_x, batch_y = next(batch_gen)
    print(batch_x, batch_y)

猜你喜欢

转载自blog.csdn.net/huanghaocs/article/details/83242353

Python产生batch数据的方法

Python循环产生批量数据batch

Tensorflow训练中产生batch

python多线程避免产生脏数据的三种方法

python中产生随机列表,随机字典,随机数据的方法

数据倾斜的产生、解决方法

Unity Batch 对 Vertex Shader 产生影响

训练数据与batch大小

spring batch元数据

Python Mocking学习笔记：产生实际数据前先造假数据

python 产生坐标的两种方法

数据库加锁死锁产生的原因和解锁的方法

Hive---数据倾斜的产生及解决方法

人工智能中噪声数据的产生与处理方法详解

python 对迭代器产生的数据进行切片

python中faker模块：产生随机数据的模块

使用python批量产生马赛克数据集

使用python批量产生水印数据集

深度学习python数据构造（二）——数据批生成器batch_generator+yield使用

TensorFlow走过的坑之---数据读取和tf中batch的使用方法

oracle 主键产生方法

Pytorch：批量数据（batch）分割

Batch

随机产生数值（Python）

python产生时间

python模拟日志产生

Go对Python产生的冲击

Python-产生密码

DataFactory 产生数据

机器学习产生数据

今日推荐

LFOSSA 源来如此公开课 | 掌握云原生未来：CNCF 认证全面攻略与备考秘籍

国产云输入法——仅华为无云端数据上传安全问题

开源日报 | 工业开源项目OGG 1.0；姐姐，你要和我一起配置火狐吗；苹果AI遥遥落后？Fedora 40

开放签电子签章：停止新增，优化体验，前进更进（五一假期前工作）

开源日报 | 中学生开源前端动画引擎；全球首个Llama3 8B中文版开源模型；联想电脑恐出局；Linus讽刺AI炒作

“百模大战”必有一战 | 2024中国“百模大战”竞争格局分析

周排行

Family Tree 题解

BZOJ 1093 最大半连通子图 SCC + DP

幂等处理

Spring----学习（2）----XML 配置Bean 自动装配

SQL Server 远程更新目标表数据

HIbernate3.6 环境搭建

特殊符号正则表达式

【Linux】第一章进程的理解

843. n-皇后问题（dfs+输出各种情况）

空间数据库2

每日归档

更多

2024-04-26(39)

2024-04-25(22)

2024-04-24(36)

2024-04-23(26)

2024-04-22(39)

2024-04-21(0)

2024-04-20(6)

2024-04-19(5)

2024-04-18(0)

2024-04-17(5)