东篱科研大数据发现系统（DRDS）

位置：成果数据库 > 期刊 > 期刊详情页

面向大数据处理的并行优化抽样聚类K-means算法

ISSN号：1001-9081
期刊名称：《计算机应用》
时间：0
分类：TP391[自动化与计算机技术—计算机应用技术;自动化与计算机技术—计算机科学与技术]
作者机构：湖南大学信息科学与工程学院,长沙410082
相关基金：国家自然科学基金资助项目（61173107）; 国家863计划项目（2012AA01A301-01）

关键词：大数据, K-均值, 概率抽样, 欧氏距离, 聚类精度, big data, K-means, probability sampling, Euclidean distance, clustering accuracy

中文摘要：

针对大数据环境下K-means聚类算法聚类精度不足和收敛速度慢的问题,提出一种基于优化抽样聚类的K-means算法（OSCK）。首先,该算法从海量数据中概率抽样多个样本;其次,基于最佳聚类中心的欧氏距离相似性原理,建模评估样本聚类结果并去除抽样聚类结果的次优解;最后,加权整合评估得到的聚类结果得到最终k个聚类中心,并将这k个聚类中心作为大数据集聚类中心。理论分析和实验结果表明,OSCK面向海量数据分析相对于对比算法具有更好的聚类精度,并且具有很强的稳健性和可扩展性。

英文摘要：

Focusing on the low accuracy and slow convergence of K-means clustering algorithm, an improved K-means algorithm based on optimization sample clustering named OSCK（ Optimization Sampling Clustering K-means Algorithm） was proposed. Firstly, multiple samples were obtained from mass data by probability sampling. Secondly, based on Euclidean distance similarity principle of optimal clustering center, the results of sample clustering were modeled and evaluated, and the sub-optimal solution of sample clustering results was removed. Finally, the final k clustering centers were got by weighted integration evaluation of clustering results, and the final k clustering centers were used as cluster centers of big data set.Theoretical analysis and experimental results show that the proposed method for mass data analysis with respect to the comparison algorithm has better clustering accuracy, and has strong robustness and scalability.

同期刊论文项目

面向动态多目标优化的量子Memetic计算策略与算法研究

期刊论文 48 会议论文 6

同项目期刊论文

A Novel Active Contour Model for Object Tracking Based on Frog Visual Characteristics

A hybrid algorithm based on particle swarm and chemical reaction optimization

Based-on SOA Security Policy for Transverse Networking System

A Novel Method for Moving Objects Detection Inspired by Frog's Visual Characteristics

Quantum-inspired Hyper-heuristics for Energy-aware Scheduling on Heterogeneous Computing Systems

Dynamic Multiobjective Optimization Algorithm Based on Average Distance Linear Prediction Model

A novel approach to delay-fractional-dependent stability criterion for linear systems with interval

A Novel Relational Database Watermarking Algorithm Based on Clustering and Polar Angle Expansion

基于生态策略的动态多目标优化算法

A hybrid algorithm based on particle swarm and chemical reaction optimization for multi-object probl

Orthogonal chemical reaction optimization algorithm for global numerical optimization problems

云环境下超启发式能耗感知调度算法

贝叶斯博弈多目标进化算法及其收敛性分析

基于生态种群捕获竞争模型的多目标Memetic优化算法

异构云环境多目标Memetic优化任务调度方法

基于正交设计的动态多目标优化算法

Proactive workload management in dynamic virtualized environments

A New Method of Motion Detection with Biological Intelligence

约束优化进化算法综述

期刊信息

《计算机应用》
北大核心期刊（2011版）

主管单位:四川省科学技术协会
主办单位:四川省计算机学会中国科学院成都分院
主编：张景中
地址：成都市人民南路四段九号科分院计算所
邮编：610041
邮箱：xzh@joca.cn
电话：028-85224283

国际标准刊号：ISSN：1001-9081
国内统一刊号：ISSN：51-1307/TP
邮发代号:62-110

获奖情况:
全国优秀科技期刊一等奖,国家期刊奖提名奖,中国期刊方阵双奖期刊,中文核心期刊,中国科技核心期刊

国内外数据库收录:
俄罗斯文摘杂志,波兰哥白尼索引,美国剑桥科学文摘,英国科学文摘数据库,日本日本科学技术振兴机构数据库,中国中国科技核心期刊,中国北大核心期刊（2004版）,中国北大核心期刊（2008版）,中国北大核心期刊（2011版）,中国北大核心期刊（2014版）,中国北大核心期刊（2000版）

被引量:53679