Class TableExtractor

java.lang.Object
org.getalp.dbnary.morphology.HtmlTableHandler
org.getalp.dbnary.morphology.TableExtractor
Direct Known Subclasses:
GermanTableExtractor, SwedishTableExtractor

public abstract class TableExtractor extends HtmlTableHandler
  • Field Details

    • alreadyParsedTables

      protected Set<org.jsoup.nodes.Element> alreadyParsedTables
    • currentEntry

      protected String currentEntry
  • Constructor Details

    • TableExtractor

      public TableExtractor()
  • Method Details

    • getInflectionDataFromCellContext

      protected abstract List<? extends InflectionData> getInflectionDataFromCellContext(List<String> context)
      returns the inflection data that correspond to current celle context

      The cell context is a list of String that corresponds to all column and row headers + section headers in which the cell appears.

      Parameters:
      context - a list of Strings that represent the celle context
      Returns:
      The InflexionData corresponding to the context
    • shouldIgnoreCurrentH2

      protected boolean shouldIgnoreCurrentH2(org.jsoup.nodes.Element elt)
      returns true if the current H2 element should be ignore while extracting morphological tables
      Parameters:
      elt - the H2 Header Element
      Returns:
      true iff the section should be ignored
    • parseHTML

      public InflectedFormSet parseHTML(String htmlCode, String pagename)
    • decodeH2Context

      protected Collection<? extends String> decodeH2Context(String text)
    • parseTable

      protected InflectedFormSet parseTable(org.jsoup.nodes.Element table, List<String> globalContext)
    • shouldProcessCell

      protected boolean shouldProcessCell(org.jsoup.nodes.Element cell)
    • handleSimpleCell

      protected void handleSimpleCell(org.jsoup.nodes.Element cell, List<String> context, InflectedFormSet forms)
    • handleNestedTables

      protected void handleNestedTables(org.jsoup.nodes.Element cell, List<String> context, InflectedFormSet forms)
    • isNormalCell

      protected boolean isNormalCell(org.jsoup.nodes.Element cell)
    • isHeaderCell

      protected boolean isHeaderCell(org.jsoup.nodes.Element cell)
    • getRowAndColumnContext

      protected List<String> getRowAndColumnContext(int nrow, int ncol, ArrayMatrix<org.jsoup.nodes.Element> columnHeaders)
    • addToContext

      protected void addToContext(ArrayMatrix<org.jsoup.nodes.Element> columnHeaders, int i, int j, List<String> res)
    • getInflectedForms

      protected Set<String> getInflectedForms(org.jsoup.nodes.Element cell)
      Extract wordforms from table cell
      Splits cell content by <br\> or comma and removes HTML formatting
      Parameters:
      cell - the current cell in the inflection table
      Returns:
      Set of wordforms (Strings) from this cell