sábado, 13 de dezembro de 2014

Randomizing with Elasticsearch: a practical example

This post explains how to return shuffle values returned by Elasticsearch. The use case is the situation where we want to avoid that users receive results of only one (or few) type. For instance, if you have a ecommerce and you want to return products of different brands, even though, you have some brands that dominates your dataset (ie., you have a brand that represents a large amount of your data).

1. In order to test that I created a small dataset with 15 products of 3 brands.

PUT /test

PUT /test/products/1
{
    "name" : "product 1",
    "brand" : "brand1"
}

PUT /test/products/2
{
    "name" : "product 2",
    "brand" : "brand1"
}

PUT /test/products/3
{
    "name" : "product 3",
    "brand" : "brand1"
}

PUT /test/products/4
{
    "name" : "product 4",
    "brand" : "brand1"
}

PUT /test/products/5
{
    "name" : "product 5",
    "brand" : "brand1"
}

PUT /test/products/6
{
    "name" : "product 6",
    "brand" : "brand2"
}

PUT /test/products/7
{
    "name" : "product 7",
    "brand" : "brand2"
}

PUT /test/products/8
{
    "name" : "product 8",
    "brand" : "brand2"
}

PUT /test/products/9
{
    "name" : "product 9",
    "brand" : "brand2"
}

PUT /test/products/10
{
    "name" : "product 10",
    "brand" : "brand2"
}


PUT /test/products/11
{
    "name" : "product 11",
    "brand" : "brand3"
}

PUT /test/products/12
{
    "name" : "product 12",
    "brand" : "brand3"
}

PUT /test/products/13
{
    "name" : "product 13",
    "brand" : "brand3"
}

PUT /test/products/14
{
    "name" : "product 14",
    "brand" : "brand3"
}

PUT /test/products/15
{
    "name" : "product 15",
    "brand" : "brand3"
}

2. I did three queries: (A) one without any sorting, (B) a sort script using the Java hashCode function and (C) the last one using Elasticsearch function of random_score.

POST /test/products/_search
{
   "from": 0,
   "size": 3,
   "query": {
      "match": {
         "name": "product"
      }
   }
}

POST /test/products/_search
{
   "from": 0,
   "size": 3,
   "query": {
      "match": {
         "name": "product"
      }
   },
   "sort": {
      "_script": {
         "script": "(doc['_id'].value + seed).hashCode()",
         "type": "number",
         "params": {
            "seed": "doc['name'].value"
         },
         "order": "asc"
      }
   }
}

POST /test/products/_search
{
   "from": 0,
   "size": 3,
   "query": {
      "function_score": {
         "query": {
            "match": {
               "name": "product"
            },
            "functions": [
               {
                  "random_score": {
                     "seed": "1"
                  }
               }
            ],
            "score_mode": "sum"
         }
      }
   }
}

3. The results were:

A. brand 2, brand 2, brand 1
B. brand 1, brand 3, brand 2
C. brand 3, brand 3, brand 1

4. My conclusion is that using the Java hash code function is the best approach. The ramdon_score function is interesting if you want to keep the results consistents for a same user (you can use the user id as the seed of this function).

Best regards,

Luiz

sexta-feira, 7 de novembro de 2014

quinta-feira, 11 de setembro de 2014

Popularity + most searched terms

In this post I will combine the popularity of a document with the most searched terms in order to boost the results based on past search issued against a particular index.

PUT blogposts

PUT statistics

PUT /blogposts/post/1
{
  "title":   "About popularity",
  "content": "In this post we will talk about...",
  "votes":   6
}

PUT /blogposts/post/2
{
  "title":   "About elasticsearch",
  "content": "In this post we will talk about...",
  "votes":   3
}

PUT /blogposts/post/3
{
  "title":   "About popularity",
  "content": "In this post we will talk about...",
  "votes":   7
}

PUT /statistics/queries/1
{
  "user_query":   "popularity"
}

PUT /statistics/queries/2
{
  "user_query":   "popularity in elasticsearch"
}

PUT /statistics/queries/3
{
  "user_query":   "boost"
}

PUT /statistics/queries/4
{
  "user_query":   "boost in elasticsearch"
}


PUT /statistics/queries/5
{
  "user_query":   "elasticsearch is the best search engine"
}

GET blogposts/post/_mapping

GET statistics/queries/_mapping

POST statistics/queries/_search
{
    "query" : {
        "match_all" : {}
    },
    "facets": {
        "keywords": {
            "terms": {
                "field": "user_query"
            }
        }
    }
}

POST blogposts/post/_search
{
 
    "sort" : [
        { "votes" : {"order" : "desc"}},
        "_score"
    ],
    "query" : {
        "match" : {
            "title":{
                "query":"elasticsearch popularity"
             
            }
        }
    }
}




quarta-feira, 23 de abril de 2014

[ElasticSearch] An example of using SetFetchSource

The SetFetchSource is a new feature of ElasticSearch1.1. This is a example of how to use it:

String [] excludes = {"field0"};
String [] includes = {"field1","field2"};

SearchResponse searcher = client.getClient().prepareSearch(INDEX).setFetchSource(includes, excludes).setQuery(qb).execute().actionGet();

If you do not include the "addFields()" command, you will not be able to iterate over the fields of the hits. However, you can iterate using "sourceAsMap()".


for (SearchHit hit : searcher.getHits().getHits()) {
    Map hits = hit.sourceAsMap();

    for (String key : hits.keySet()) {
        //do something
    }

    Object [] fieldValues = hits.values().toArray();

    for (Object fieldValue:fieldValues) {
        //do something
    }
}

As you will notice, the excluded fields at "excludes" will not show up. Also, if you try "hit.getFields()" it will be empty.



References:

http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/mapping-source-field.html

https://github.com/elasticsearch/elasticsearch/blob/master/src/test/java/org/elasticsearch/search/source/SourceFetchingTests.java


terça-feira, 22 de abril de 2014

[ElasticSearch] QueryBuild vs. String: ElasticsearchParseException[Failed to derive xcontent from...

I had a problem while writing tests for use the new setFetchSource feature of ES 1.1. After some test, research and thinking (try and error). I finally realised that the problem rises when you are using a plain string query instead of a QueryBuilder.


org.elasticsearch.action.search.SearchPhaseExecutionException: Failed to execute phase [query_fetch], all shards failed; shardFailures {[KMpFYBpxRECgm3gJC_1-uw][content][0]: SearchParseException[[content][0]: from[-1],size[10]: Parse Failure [Failed to parse source [{"size":10,"query_binary":"VGVzdHRleHQ=","_source":{"includes":["CONTENTID","URL"],"excludes":["CONTENT"]},"fields":["CONTENTID","URL"]}]]]; nested: ElasticsearchParseException[Failed to derive xcontent from (offset=0, length=8): [84, 101, 115, 116, 116, 101, 120, 116]]; }
    at org.elasticsearch.action.search.type.TransportSearchTypeAction$BaseAsyncAction.onFirstPhaseResult(TransportSearchTypeAction.java:272)
    at org.elasticsearch.action.search.type.TransportSearchTypeAction$BaseAsyncAction$3.onFailure(TransportSearchTypeAction.java:224)
    at org.elasticsearch.search.action.SearchServiceTransportAction.sendExecuteFetch(SearchServiceTransportAction.java:307)
    at org.elasticsearch.action.search.type.TransportSearchQueryAndFetchAction$AsyncAction.sendExecuteFirstPhase(TransportSearchQueryAndFetchAction.java:71)
    at org.elasticsearch.action.search.type.TransportSearchTypeAction$BaseAsyncAction.performFirstPhase(TransportSearchTypeAction.java:216)
    at org.elasticsearch.action.search.type.TransportSearchTypeAction$BaseAsyncAction.performFirstPhase(TransportSearchTypeAction.java:203)
    at org.elasticsearch.action.search.type.TransportSearchTypeAction$BaseAsyncAction$2.run(TransportSearchTypeAction.java:186)
    at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)
    at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)
    at java.lang.Thread.run(Thread.java:744)


Makes sense, since the error says: "failed to derive xcontent from (offset". However, for a not native english speaker, took some time. Just change:

String query="";
        
SearchResponse response = client.prepareSearch(CONTENT_INDEX).setFetchSource(includes, excludes).setQuery(query).addFields(includes)
                    .setSize(10).execute().actionGet();


To:

String query="";

QueryStringQueryBuilder qb = QueryBuilders.queryString(query);

SearchResponse response = client.prepareSearch(CONTENT_INDEX).setFetchSource(includes, excludes).setQuery(qb).addFields(includes)
                    .setSize(10).execute().actionGet();

terça-feira, 15 de abril de 2014

SecurityException: WifiService: Neither user 10082 nor current process has android.permission.ACCESS_WIFI_STATE

If you are trying to get scan the networks through a mobile device, first of all you have to add the permissions in your AndroidManifest.xml:








public class MainActivity ... {

 public void scanWiFi(View view) {
        mainWifi = (WifiManager) getSystemService(Context.WIFI_SERVICE);

        registerReceiver(receiverWifi, new IntentFilter(WifiManager.SCAN_RESULTS_AVAILABLE_ACTION));
        mainWifi.startScan();

}

    class WifiReceiver extends BroadcastReceiver {
        public void onReceive(Context c, Intent intent) {
            StringBuilder sb = new StringBuilder();
            wifiList = mainWifi.getScanResults();
            for (int i = 0; i < wifiList.size(); i++) {
                sb.append(new Integer(i + 1).toString() + ".");
                sb.append((wifiList.get(i)).SSID.toString());
                sb.append((wifiList.get(i)).frequency);
                sb.append((wifiList.get(i)).level);
                sb.append("\\n");
            }
            mainText.append(sb);
        }
    }

}

terça-feira, 8 de abril de 2014

How to use the Path Hierarchy Tokenizer in ElasticSearch

How to use the Path Hierarchy Tokenizer:

0. If you and just to check the output of this tokenizer, you can run:

curl -XGET 'localhost:9200/_analyze?tokenizer=path_hierarchy&filters=lowercase' -d '/something/something/else'


1. Put a map:

$ curl -XPUT localhost:9200/index_files3/ -d '
{
  "settings":{
     "index":{
        "analysis":{
           "analyzer":{
              "analyzer_path_hierarchy":{
                 "tokenizer":"my_path_hierarchy_tokenizer",
                 "filter":"lowercase"
              }
           },
       "tokenizer" : {
              "my_path_hierarchy_tokenizer" : {
                  "type" : "path_hierarchy",
                  "delimiter" : "/",
                  "replacement" : "*",
                  "buffer_size" : "1024",
                  "reverse" : "true",
                  "skip" : "0"
               }
           }
        }
     }
  },
  "mappings":{
     "file":{
        "properties":{
           "path":{
              "analyzer":"analyzer_path_hierarchy",
              "type":"string"
           }
        }
     }
  }
}'


2. Get the map:

curl -XGET 'http://localhost:9200/index_files3/file/_mapping'


3. Add a new document:

$ curl -XPUT 'http://localhost:9200/index_files3/file/1' -d '{
    "name" : "c1",
    "text" : "t1",
    "path" : "/c1/c2/c3"
}'


4. Check if it is there:

$ curl -XGET 'http://localhost:9200/index_files3/file/1'


5. Search:

Fail: curl -XGET 'http://localhost:9200/index_files3/_search?q=path:c1'

Success: curl -XGET 'http://localhost:9200/index_files3/_search?q=path://c1'