首页 | 本学科首页   官方微博 | 高级检索  
     

三元家庭基因数据的单体分型和单体型频率估计
引用本文:张强锋,徐云,陈国良,车皓阳. 三元家庭基因数据的单体分型和单体型频率估计[J]. 软件学报, 2007, 18(9): 2090-2099
作者姓名:张强锋  徐云  陈国良  车皓阳
作者单位:中国科学技术大学,计算机系,安徽,合肥,230027;中国科学院,软件研究所,计算机科学国家重点实验室,北京,100080;中国科学技术大学,计算机系,安徽,合肥,230027;中国科学院,软件研究所,计算机科学国家重点实验室,北京,100080
摘    要:研究了在门德尔遗传定理和哈代-维恩伯格平衡假设下,三元家庭基因型数据的单体分型和单体型频率估计问题.过去的研究仅仅关注个体间没有联系或者含有一般家系信息的基因型数据,而对这种特殊的三元家庭关注得不够.考虑到HAPMAP数据库中有一部分数据就基于这种三元家庭,现在有越来越多的需求要求直接分析这种特殊的家系结构.提出一个两段式的三元家庭中单体型频率的估计方法:i) 分型阶段,找出每一个三元家庭零重组单体构型;ii) 频率估计阶段,在前一阶段得到的单体构型基础上,应用EM算法来估计单体型频率.在程序包TRIOHAP中用C语言实现了单体分型算法和EM算法,并且使用模拟和实际数据测试了TRIOHAP的有效性和效率.实验结果表明,TRIOHAP要比其他那些忽略了三元家庭信息的常见单体型频率估计软件运行快很多.进一步地,由于TRIOHAP利用了这些信息,其估计结果更加可靠.

关 键 词:基因型  单体型  SNP  单体分型  单体型频率估计  三元家庭  EM算法
收稿时间:2004-12-21
修稿时间:2004-12-212006-03-31

Haplotyping and Haplotype Frequency Estimates on Trio Genotype Data
ZHANG Qiang-Feng,XU Yun,CHEN Guo-Liang and CHE Hao-Yang. Haplotyping and Haplotype Frequency Estimates on Trio Genotype Data[J]. Journal of Software, 2007, 18(9): 2090-2099
Authors:ZHANG Qiang-Feng  XU Yun  CHEN Guo-Liang  CHE Hao-Yang
Affiliation:1.Department of Computer Science, University of Science and Technology of China, Hefei 230027, China;State Key Laboratory of Computer Science, Institute of Software, The Chinese Academy of Sciences, Beijing 100080, China
Abstract:The problems of haplotyping and haplotype frequency estimation on trio genotype data under the Mendelian law of inheritance and the assumption of Hardy-Weinberg equilibrium are studied in this paper. Since most past efforts only focused on haplotyping on genotype data of unrelated individuals and data with general pedigrees, but gave insufficient efforts to the special case of trio genotype data, there is coming an increasing demand in analyzing them in particular, especially when taking into account that part of HAPMAP database is exactly trio data. This paper presents a two-staged method to estimate haplotype frequencies in trios: i) haplotyping stage, find haplotype configurations without recombinant for each trio; ii) frequency estimation stage, use the expectation-maximization (EM) algorithm to estimate haplotype frequencies based on these inferred haplotype configurations. Both the haplotyping algorithm and the EM algorithm are implemented in software package TRIOHAP using C language. Its effectiveness and efficiency and tested on simulated and real data sets as well. The experimental results show that, TRIOHAP runs much faster than a popular frequency estimation software which discards trio information. Moreover, because TRIOHAP utilizes such information, its estimation is more reliable.
Keywords:SNP  genotype  haplotype  SNP  haplotyping  haplotype frequencies estimate  trio  EM algorithm
本文献已被 CNKI 维普 万方数据 等数据库收录!
点击此处可从《软件学报》浏览原始摘要信息
点击此处可从《软件学报》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号