WANG Ru, SONG Han-tao, LU Yu-chang. Web Pages Data Extraction Based on Tree AutomataJ. Transactions of Beijing institute of Technology, 2004, (9): 790-793.
Citation: WANG Ru, SONG Han-tao, LU Yu-chang. Web Pages Data Extraction Based on Tree AutomataJ. Transactions of Beijing institute of Technology, 2004, (9): 790-793.

Web Pages Data Extraction Based on Tree Automata

  • In order to extract data from HTML Web pages automatically, tree automata induction has been used in data extraction. The key idea is to transform the example tree into a binary tree, creating a tree automata which can accept the binary tree of example pages and using the tree automata to extract data according to tree automata state of acceptance and rejection. The method makes use of the native tree structure of HTML document and designs a new simple form of labeling the example pages. Experimental results on data sets showed that the approach with tree automata compared favorable against some other approaches in the F-score and recall.
  • loading

Catalog

    Turn off MathJax
    Article Contents

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return
    Baidu
    map