暂无图片
暂无图片
暂无图片
暂无图片
暂无图片
基于弱匹配概率典型相关性分析的图像自动标注-张博 , 郝杰 , 马刚 , 史忠植.pdf
59
18页
1次
2022-05-20
免费下载
软件学报 ISSN 1000-9825, CODEN RUXUEW E-mail: jos@iscas.ac.cn
Journal of Software,2017,28(2):292309 [doi: 10.13328/j.cnki.jos.005047] http://www.jos.org.cn
©中国科学院软件研究所版权所有. Tel: +86-10-62562563
基于弱匹配概率典型相关性分析的图像自动标注
1
,
4
,
2,3
,
史忠植
2
1
(中国矿业大学 计算机科学与技术学院,江苏 徐州 221116)
2
(中国科学院 计算技术研究所 智能信息处理重点实验室,北京 100190)
3
(中国科学院大,北京 100049)
4
(徐州医科大学 医学信息学院,江苏 徐州 221004)
通信作者: 郝杰, E-mail: haojie@xzmc.edu.cn
: 针对弱匹配多模态数据的相关性建模问题,提出了一种弱匹配概率典型相关性分析模型(semi-paired
probabilistic CCA,简称 SemiPCCA).SemiPCCA 模型关注于各模态内部的全局结构,模型参数的估计受到了未匹配
样本的影响,未匹配样本则揭示了各模态样本空间的全局结构.在人工弱匹配多模态数据集上的实验结果表
,SemiPCCA 可以有效地解决传统 CCA(canonical correlation analysis) PCCA(probabilistic CCA)在匹配样本不足
的情况下出现的过拟合问题,取得了较好的效果.提出了一种基于 SemiPCCA 的图像自动标注方法.该方法基于关联
建模的思想,同时使用标注图像及其关键词和未标注图像学习视觉模态和文本模态之间的关联,从而能够更准确地
对未知图像进行标注.
关键词: 典型相关性分析;概率典型相关性分析;弱匹配典型相关性分析;图像自动标注
中图法分类号: TP391
中文引用格式: 张博,郝杰,马刚,史忠植.基于弱匹配概率典型相关性分析的图像自动标注.软件学报,2017,28(2):292–309.
http://www.jos.org.cn/1000-9825/5047.htm
英文引用格式: Zhang B, Hao J, Ma G, Shi ZZ. Automatic image annotation based on semi-paired probabilistic canonical
correlation analysis. Ruan Jian Xue Bao/Journal of Software, 2017,28(2):292309 (in Chinese). http://www.jos.org.cn/1000-9825/
5047.htm
Automatic Image Annotation Based on Semi-Paired Probabilistic Canonical Correlation
Analysis
ZHANG Bo
1
, HAO Jie
4
, MA Gang
2,3
, SHI Zhong-Zhi
2
1
(School of Computer Science and Technology, China University of Mining and Technology, Xuzhou 221116, China)
2
(Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, The Chinese Academy of Sciences, Beijing
100190, China)
3
(University of Chinese Academy of Sciences, Beijing 100049, China)
4
(School of Medicine Information, Xuzhou Medical University, Xuzhou 221004, China)
Abstract: Canonical correlation analysis (CCA) is a statistical analysis tool for analyzing the correlation between two sets of random
variables. CCA requires the data be rigorously paired or one-to-one correspondence among different views due to its correlation definition.
However, such requirement is usually not satisfied in real-world applications due to various reasons. Often, only a few paired and a lot of
基金项目: 国家重点基础研究发展计划(973)(2013CB329502); 国家自然科学基金(61035003); 国家高技术研究发展计划(863)
(2012AA011003); 国家科技支撑计划(2012BA107B02); 江苏省自然科学基金(BK20160276)
Foundation item: National Program on Key Basic Research Project of China (973) (2013CB329502); National Natural Science
Foundation of China (61035003); National High-Tech R&D Program of China (863) (2012AA011003); National Key Technology R&D
Program of China (2012BA107B02); Natural Science Foundation of Jiangsu Province (BK20160276)
收稿时间: 2014-12-18; 修改时间: 2015-06-11, 2015-09-10; 采用时间: 2016-02-03
张博 :基于弱匹配概率典型相关性分析的图像自动标注
293
unpaired multi-view data are given, because unpaired multi-view data are relatively easier to be collected and pairing them is difficult,
time consuming and even expensive. Such data is referred as semi-paired multi-view data. When facing semi-paired multi-view data, CCA
usually performs poorly. To tackle this problem, a semi-paired variant of CCA, named SemiPCCA, is proposed based on the probabilistic
model for CCA. The actual meaning of “semi-” in SemiPCCA is “semi-paired” rather than “semi-supervised” as in popular
semi-supervised learning literature. The estimation of SemiPCCA model parameters is affected by the unpaired multi-view data which
reveal the global structure within each modality. By using artificially generated semi-paired multi-view data sets, the experiment shows
that SemiPCCA effectively overcome the over-fitting problem of traditional CCA and PCCA (probabilistic CCA) under the condition of
insufficient paired multi-view data and performs better than the original CCA and PCCA. In addition, an automatic image annotation
method based on the SemiPCCA is presented. Through estimating the relevance between images and words by using the labelled and
unlabeled images together, this method is shown to be more accurate than previous published methods.
Key words: canonical correlation analysis; probabilistic canonical correlation analysis; semi-paired canonical correlation analysis;
automatic image annotation
物联网、互联网等拥有丰富的文本、图像、视频和音频等多媒体信息资源,这些信息资源是异构的,很难
直接发现它们之间的关联.目前,典型相关性分析(canonical correlation analysis,简称 CCA)作为一种分析两组随
机变量之间相关性的统计分析工具,已被引入跨媒体的相关性建模中,挖掘不同模态内容特征之间潜在的统计
相关性
[1,2]
.通过特征子空间映射,将各模态的数据从原始高维特征空间映射到低维特征空间,既解决了不同
型数据间的异构性问题,消除了多模态数据间的内容鸿,最大程度地保持了初始的相关性不变,将不同类型的
多媒体数据在特征层面上关联起来,同时也最大程度地保持初始的相关性不变.
典型相关性分析中两组相关的随机变量可以来自多种信息来源(如同一个人的声音和图像),也可以是从同
一来源的信息中抽取的不同特(如图像的颜色特征和纹理特征),但训练数据必须一对一严格匹配.很多原因
造成这种严格匹配的训练数据难以获得,:(1) 多传感器采集系统中传感器采样频率不同步或传感器故障,
造成不同通道采集来的数据不同步或丢失某一通道数据;(2) 单模态数据比较容易获得,但人工匹配却非常费
时、费力.实际中,我们面对的多模态数据经常是只有少量一对一严格匹配,其余大量数据未匹配.我们称其为弱
匹配多模态数据.
面向弱匹配多模态数据的典型相关性分析有两种基本方法:(1) 丢弃未匹配数据,只使用典型相关性分析
处理严格匹配的多模态数据;(2) 根据特定准则,匹配多模态数据.但这两种方法都不可能获得理想的结果.
本文的主要工作包括:(1) 提出了一种全新的弱匹配概率典型相关性分析模型(semi-paired probabilistic
CCA,简称 SemiPCCA).不同于以往的弱匹配典型相关性分析模型,SemiPCCA 完全基于概率典型相关性分析模
(probabilistic CCA,简称 PCCA),关注于各模态内部的全局结构,模型参数的估计受到了未匹配样本的影响,
未匹配样本则揭示了各领域样本空间的全局结构.(2) 提出了一种基于 SemiPCCA 的图像自动标注方法.该方
法同时使用标注图像及其关键词和未标注图像估计隐空间的分布,学习视觉模态和文本模态之间的关联,能够
较好地对未知图像进行标注.
1 相关工作
1.1 典型相关性分析
传统的特征分析方法, PCA(principal component analysis),ICA(independent component analysis) PLS(partial
least squares),大多用于单模态的特征分析,实现主成分提取、去噪、维数约减和保持本征度量等目的,不能同时分
析不同类型的异构特征,难以发现多种特征间的关联信息.典型相关性分析(canonical correlation analysis,CCA)是一
种用来分析两组随机变量之间相关性的统计分析工具,其相关性保持特征己经在理论上得到证明,应用于经济学
气象和基因组数据分析等领域.CCA 通过统计方法找到两组异构多模态特征之间的潜在关系,从底层特征上用统
一的模型将不同类型的多模态数据关联起来,同时尽可能地发现和保持数据间潜在的相关性.
维度分别为 p q 的两组随机变量 x y,给定均值为 0 的成对观察样本集合
{}
1
(,) ,
n
i
i
p
q
i
R
R
=
∈×xy
of 18
免费下载
【版权声明】本文为墨天轮用户原创内容,转载时必须标注文档的来源(墨天轮),文档链接,文档作者等基本信息,否则作者和墨天轮有权追究责任。如果您发现墨天轮中有涉嫌抄袭或者侵权的内容,欢迎发送邮件至:contact@modb.pro进行举报,并提供相关证据,一经查实,墨天轮将立刻删除相关内容。

评论

关注
最新上传
暂无内容,敬请期待...
下载排行榜
Top250 周榜 月榜