首页 | 本学科首页   官方微博 | 高级检索  
     

基于文本聚类的语言韵律和节奏风格特征挖掘
引用本文:贺湘情,刘颖. 基于文本聚类的语言韵律和节奏风格特征挖掘[J]. 中文信息学报, 2014, 28(6): 194-200
作者姓名:贺湘情  刘颖
作者单位:清华大学 人文学院中国语言文学系,北京 100084
基金项目:国家自然科学基金(61171114);教育部自主科研项目(20111081010)
摘    要:该文以朱自清、汪曾祺和刘亮程的散文作品为语料,旨在从文本的韵律和节奏出发,采用文本聚类的方法来挖掘出新的能够代表作品风格的特征。实验表明,以句末用字韵母的n元组合、分句句长的n元组合、标点符号和整句句长作为风格特征,能成功地将这三位作家的作品区分开来。其中刘亮程句尾韵的舌位高于汪、朱二人,朱自清对韵脚的选择不如刘、汪二人丰富。汪曾祺的分句长最短,且最为讲究句式长短的对齐;刘亮程兼顾长短句的交错,节奏更富于变化;朱自清的句长变化最为平稳。

关 键 词:特征挖掘  韵律  节奏  文本聚类  

Mining Stylistic Features of Rhythm and Tempo Based on Text Clustering
HE Xiangqing,LIU Ying. Mining Stylistic Features of Rhythm and Tempo Based on Text Clustering[J]. Journal of Chinese Information Processing, 2014, 28(6): 194-200
Authors:HE Xiangqing  LIU Ying
Affiliation:Department of Chinese Language and Literature, School of Humanity, Tsinghua University, Beijing 100084, China
Abstract:We selected literary proses written by Ziqing Zhu, Zengqi Wang and Liangcheng Liu as corpora. Text clustering is used to mine new stylistic features from the perspective of rhythm and tempo. The experimental results show that n-grams based on the vowels of the last character of the sentence, n-grams based on the length of clauses, punctuations and length of sentences, all can successfully distinguish from the articles of the three authors. Specifically, Liangcheng Liu preferred to utilize the vowels of higher tongue position. Ziqing Zhu focused on some specific rhymes, but the rhymes used by Liu and Wang are more plentiful than those of Zhu. Wang’s Clauses are the shortest, and he paid more attention to the order of sentence patterns. Long sentences and short sentences are alternatively used by Liu, and the tempos used by Liu are changeful. The sentence lengths used by Zhu are less changeful.
Keywords:Feature Mining   Rhythm   Tempo   Text Clustering  
本文献已被 CNKI 等数据库收录!
点击此处可从《中文信息学报》浏览原始摘要信息
点击此处可从《中文信息学报》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号