nori_readingform token filter
The nori_readingform
token filter rewrites tokens written in Hanja to their Hangul form.
PUT nori_sample
{
"settings": {
"index": {
"analysis": {
"analyzer": {
"my_analyzer": {
"tokenizer": "nori_tokenizer",
"filter": [ "nori_readingform" ]
}
}
}
}
}
}
GET nori_sample/_analyze
{
"analyzer": "my_analyzer",
"text": "鄕歌" 1
}
- A token written in Hanja: Hyangga
Which responds with:
{
"tokens" : [ {
"token" : "향가", 1
"start_offset" : 0,
"end_offset" : 2,
"type" : "word",
"position" : 0
}]
}
- The Hanja form is replaced by the Hangul translation.