056_Sorting/88_String_sorting.asciidoc · 缠中说禅/elasticsearch-definitive-guide - Gitee.com

加入 Gitee

与超过 1200万开发者一起发现、参与优秀开源项目，私有仓库也完全免费：）

文件

克隆/下载

88_String_sorting.asciidoc 2.88 KB

[[multi-fields]]
=== String Sorting and Multifields

Analyzed string fields are also multivalue fields,((("strings", "sorting on string fields")))((("analyzed fields", "string fields")))((("sorting", "string sorting and multifields"))) but sorting on them seldom
gives you the results you want. If you analyze a string like `fine old art`,
it results in three terms. We probably want to sort alphabetically on the
first term, then the second term, and so forth, but Elasticsearch doesn't have this
information at its disposal at sort time.

You could use the `min` and `max` sort modes (it uses `min` by default), but
that will result in sorting on either `art` or `old`, neither of which was the
intent.

In order to sort on a string field, that field should contain one term only:
the whole `not_analyzed` string.((("not_analyzed string fields", "sorting on")))  But of course we still need the field to be
`analyzed` in order to be able to query it as full text.

The naive approach to indexing the same string in two ways would be to include
two separate fields in the document: one that is  `analyzed` for searching,
and one that is `not_analyzed` for sorting.

But  storing the same string twice in the `_source` field is waste of space.
What we really want to do is to pass in a _single field_ but to _index it in two different ways_. All of the _core_ field types (strings, numbers,
Booleans, dates) accept a `fields` parameter ((("mapping (types)", "transforming simple mapping to multifield mapping")))((("types", "core simple field types", "accepting fields parameter")))((("fields parameter")))((("multifield mapping")))that allows you to transform a
simple mapping like:

[source,js]
--------------------------------------------------
"tweet": {
    "type":     "string",
    "analyzer": "english"
}
--------------------------------------------------

into a _multifield_ mapping like this:

[source,js]
--------------------------------------------------
"tweet": { <1>
    "type":     "string",
    "analyzer": "english",
    "fields": {
        "raw": { <2>
            "type":  "string",
            "index": "not_analyzed"
        }
    }
}
--------------------------------------------------
// SENSE: 056_Sorting/88_Multifield.json

<1> The main `tweet` field is just the same as before: an `analyzed` full-text
    field.
<2> The new `tweet.raw` subfield is `not_analyzed`.

Now, or at least as soon as we have reindexed our data, we can use the `tweet`
field for search and the `tweet.raw` field for sorting:

[source,js]
--------------------------------------------------
GET /_search
{
    "query": {
        "match": {
            "tweet": "elasticsearch"
        }
    },
    "sort": "tweet.raw"
}
--------------------------------------------------
// SENSE: 056_Sorting/88_Multifield.json

WARNING: Sorting on a full-text `analyzed` field can use a lot of memory.  See
<<aggregations-and-analysis>> for more information.

一键复制编辑原始数据按行查看历史

提交于 2016-05-31 22:46 . Colon added before code snippet (#516)

String Sorting and Multifields

Analyzed string fields are also multivalue fields, but sorting on them seldom gives you the results you want. If you analyze a string like fine old art, it results in three terms. We probably want to sort alphabetically on the first term, then the second term, and so forth, but Elasticsearch doesn’t have this information at its disposal at sort time.

You could use the min and max sort modes (it uses min by default), but that will result in sorting on either art or old, neither of which was the intent.

In order to sort on a string field, that field should contain one term only: the whole not_analyzed string. But of course we still need the field to be analyzed in order to be able to query it as full text.

The naive approach to indexing the same string in two ways would be to include two separate fields in the document: one that is analyzed for searching, and one that is not_analyzed for sorting.

But storing the same string twice in the source field is waste of space. What we really want to do is to pass in a _single field but to index it in two different ways. All of the core field types (strings, numbers, Booleans, dates) accept a fields parameter that allows you to transform a simple mapping like:

"tweet": {
    "type":     "string",
    "analyzer": "english"
}

into a multifield mapping like this:

"tweet": { (1)
    "type":     "string",
    "analyzer": "english",
    "fields": {
        "raw": { (2)
            "type":  "string",
            "index": "not_analyzed"
        }
    }
}

The main tweet field is just the same as before: an analyzed full-text field.
The new tweet.raw subfield is not_analyzed.

Now, or at least as soon as we have reindexed our data, we can use the tweet field for search and the tweet.raw field for sorting:

GET /_search
{
    "query": {
        "match": {
            "tweet": "elasticsearch"
        }
    },
    "sort": "tweet.raw"
}

Warning

Sorting on a full-text analyzed field can use a lot of memory. See [aggregations-and-analysis] for more information.

Loading...

马建仓 AI 助手

尝试更多

代码解读

代码找茬

代码优化

1

https://gitee.com/SFAC_hds/elasticsearch-definitive-guide.git

[email protected]:SFAC_hds/elasticsearch-definitive-guide.git

SFAC_hds

elasticsearch-definitive-guide

elasticsearch-definitive-guide

master