US 12,393,596 B2
Automated sampling of query results for training of a query engine
Nicholas Cooley, Oakland, CA (US)
Assigned to Maplebear Inc., San Francisco, CA (US)
Filed by Maplebear Inc., San Francisco, CA (US)
Filed on Feb. 28, 2024, as Appl. No. 18/590,055.
Application 18/590,055 is a continuation of application No. 17/826,162, filed on May 27, 2022, granted, now 11,947,551.
Prior Publication US 2024/0202204 A1, Jun. 20, 2024
This patent is subject to a terminal disclaimer.
Int. Cl. G06F 16/24 (2019.01); G06F 16/2453 (2019.01); G06F 16/2455 (2019.01); G06F 16/2457 (2019.01)
CPC G06F 16/24578 (2019.01) [G06F 16/24542 (2019.01); G06F 16/2455 (2019.01)] 20 Claims
OG exemplary drawing
 
1. A method, comprising:
at a computer system comprising at least one processor and memory:
receiving a plurality of search queries from a plurality of users directed at an item query engine of an online system, each search query including a search phrase used by a user to conduct the search query;
generating a training set comprising a plurality of training samples, the plurality of training samples comprising a plurality of search phrases, each training sample corresponding to a search query of the item query engine and comprising a search phrase and a list of items returned by the item query engine;
generating a representative set of training samples, wherein generating the representative set of training samples comprises:
determining search frequencies of the search phrases used in the training samples,
stratifying the set of training samples into a plurality of bins according to the search frequencies of the search phrases, wherein each bin of the plurality of bins includes a subset of training samples, wherein each bin defines a range of numbers of times a search phrase is used by the plurality of users, and
sampling the training samples from the plurality of bins to collect the representative set of training samples; and
training the item query engine using the representative set of training samples.