Class AbstractWiktionaryExtractor
java.lang.Object
org.getalp.dbnary.languages.AbstractWiktionaryExtractor
- All Implemented Interfaces:
IWiktionaryExtractor
- Direct Known Subclasses:
FunctionalWiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor, WiktionaryExtractor
-
Field Summary
FieldsModifier and TypeFieldDescriptionprotected static final Stringprotected ExpandAllWikiModelprotected Stringprotected IWiktionaryDataHandlerprotected WiktionaryPageSourceprotected Stringprotected static final Pattern -
Constructor Summary
Constructors -
Method Summary
Modifier and TypeMethodDescriptionstatic StringcleanUpMarkup(String group) static StringcleanUpMarkup(String str, boolean humanReadable) cleans up the wiktionary markup from a string in the following manner:
str is the string to be cleaned up.protected intcomputeRegionEnd(int blockStart, Matcher m) voidcomputeStatistics(String dumpVersion) static Stringabstract voidvoidextractData(String wiktionaryPageName, String pageContent) org.apache.jena.rdf.model.ResourceextractDefinition(String definition, int defLevel) Extract and register a new wordsense in the current lexical entry.org.apache.jena.rdf.model.ResourceextractDefinition(Matcher definitionMatcher) Extract and register a new wordsense in the current lexical entry.protected voidextractDefinitions(int startOffset, int endOffset) org.apache.jena.rdf.model.ResourceextractExample(String example) voidextractExample(Matcher definitionMatcher) protected voidextractNyms(String synRelation, int startOffset, int endOffset) protected voidextractOrthoAlt(int startOffset, int endOffset) booleanfilterOutPage(String pagename) protected StringvoidpopulateMetadata(String dumpFilename, String extractorVersion) voidpostProcessData(String dumpFileVersion) voidpostProcessModel(org.apache.jena.rdf.model.Model enhancementModel, org.apache.jena.rdf.model.Model sourceModel, String dumpFileVersion) static Stringvoidprotected voidsetWiktionaryPageName(String wiktionaryPageName) static Stringprotected StringvalidateAndStandardizeLanguageCode(String language) Standardize a wiktionary language code into a "valid" language code.
-
Field Details
-
pageContent
-
wdh
-
expander
-
wiktionaryPageName
-
wi
-
debutOrfinDecomPatternString
-
xmlCommentPattern
-
NON_STANDARD_LANGUAGE_MAPPINGS
-
-
Constructor Details
-
AbstractWiktionaryExtractor
-
-
Method Details
-
setWiktionaryIndex
- Specified by:
setWiktionaryIndexin interfaceIWiktionaryExtractor
-
getWiktionaryPageName
-
setWiktionaryPageName
-
removeXMLComments
-
extractData
- Specified by:
extractDatain interfaceIWiktionaryExtractor
-
filterOutPage
- Parameters:
pagename- the name of the page- Returns:
- returns true iff the pagename should be ignored during extraction.
-
extractData
public abstract void extractData() -
extractDefinitions
protected void extractDefinitions(int startOffset, int endOffset) -
extractDefinition
Extract and register a new wordsense in the current lexical entry. The definition is extracted from a Matcher object where the group contains the list item's content and group(1) the definition. The definition level is taken from the matcher object- Parameters:
definitionMatcher- the pattern matcher containing the definition- Returns:
- the resource representing the new word sense
-
extractDefinition
Extract and register a new wordsense in the current lexical entry.- Parameters:
definition- the definition stringdefLevel- the level at which the word sense is defined (sub-senses vs senses)- Returns:
- the resource representing the new word sense
-
cleanUpMarkup
-
extractExample
-
validateAndStandardizeLanguageCode
Standardize a wiktionary language code into a "valid" language code. As language editions use codes that may differ from ISO-639-3. Sometimes these codes are referring to languages that are not represented in the iso standard and there are some that may lead to invalid turtle dumps (either because IRI become malformed, but also because the language tag of string values is invalid.In this common implementation, we only consider ISO language codes.
Language extractor may refine this method or just add new language to the NON_STANDARD_LANGUAGE_MAPPINGS map.
- Parameters:
language- the language code to be checked- Returns:
- the String representing the standardized representation for the language (usable as a language tag in RDF) or null if language is invalid
-
extractExample
-
cleanUpMarkup
cleans up the wiktionary markup from a string in the following manner:
str is the string to be cleaned up. the result depends on the value of humanReadable. Wiktionary macros are always discarded. xml/xhtml comments are always discarded. Wiktionary links are modified depending on the value of humanReadable. e.g. str = "{{a Macro}} will be [[discard]]ed and [[feed|fed]] to the [[void]]." if humanReadable is true, it will produce: "will be discarded and fed to the void." if humanReadable is false, it will produce: "will be #{discard|discarded}# and #{feed|fed}# to the #{void|void}#."- Parameters:
str- is the String to be cleaned uphumanReadable- a boolean- Returns:
- a String
-
convertToHumanReadableForm
-
extractOrthoAlt
protected void extractOrthoAlt(int startOffset, int endOffset) -
computeRegionEnd
-
extractNyms
-
stripParentheses
-
postProcessData
- Specified by:
postProcessDatain interfaceIWiktionaryExtractor
-
postProcessModel
public void postProcessModel(org.apache.jena.rdf.model.Model enhancementModel, org.apache.jena.rdf.model.Model sourceModel, String dumpFileVersion) -
computeStatistics
- Specified by:
computeStatisticsin interfaceIWiktionaryExtractor
-
populateMetadata
- Specified by:
populateMetadatain interfaceIWiktionaryExtractor
-