一种中文领域网页过滤方法

A Method of Filtering Chinese Webpage

  • 摘要: 鉴于互联网上各种不良网页的影响,提出了一种使用贝叶斯分类算法和领域本体过滤中文网页的方法。 该方法根据正反例领域网页计算领域特征词的权重,建立领域特征词库并制作领域本体,根据正例领域网页得到本体元素权重库;使用贝叶斯分类算法得到候选网页;根据领域本体对候选网页进行语义相关度计算并进行网页过滤。 该方法可以区分相同领域网页中的正反例网页并可兼顾网页过滤的实时性。 通过游戏领域网页的测试,准确率和召回率均在98%以上, 语义分析游戏相关网页的平均时间为1~2 s, 对用户浏览网页速度的影响较小, 效果令人满意。

     

    Abstract: In view of the adverse effects of a variety of useless webpages, a method based on the Bayesian classification algorithm and domain ontology was proposed to filter the unwanted Chinese webpages. The method firstly calculated the weight of domain feature words according to the positive and negative domain webpages, established domain feature lexicon and constructed the domain ontology, got the weights library of ontology elements according to the positive domain webpages; then acquired the candidates by using the Bayesian classification algorithm; lastly semantically analyzed and filtered the candidates according to the domain ontology. This method can not only distinguish the positive and negative webpages which are in the same field but also get a good performance on the real-time of webpages filtering. The experiments on huge numbers of game-related webpages have shown promising results. The precision and recall are more than 98%, the average time of semantically analyzing one game webpage is 1~2 s, it has little effect on user browsing webpages.

     

/

返回文章
返回
Baidu
map