Hi,
Some of you may be familiar with http://commoncrawl.org ; they are doing an excellent job of making large crawls of the web accessible to everyone.
I've been working on an open search engine based on these crawls for a while, and I would love to have your feedbacks on the project: https://about.commonsearch.org/
Specifically, I would be curious to know what you would consider to be the best possible integration of Wikipedia & Wikidata in a general search engine?
As a first step, we have just started using the "official website" property from Wikidata and we are considering importing the Wikipedia abstracts next (https://github.com/commonsearch/cosr-back/issues/11).
I'm looking forward to your feedbacks... or contributions! :-)
Thanks in advance,
PS: A few wikimedians recommended me to post on wikitech-l to keep the focus on the technical aspects of the project and hopefully avoid linking this project in any way to the KE stuff, which it actually predates by far (https://news.ycombinator.com/item?id=6209088).
-- Sylvain Zimmer http://sylvinus.org
Greetings Sylvian,
The Discovery department has been talking pretty actively with WMDE about this very topic. We've already explored WikiData descriptions in search results, brainstormed about how relevance functions could be affected by WikiData, seen great community tools like ppp-sparql connect natural language search and WikiData through WDQS [1], and started to think about structured data on commons [2].
If any of those interest you, then please join us on irc #wikimedia-discovery and we can chat more.
--tomasz
[1] - https://tools.wmflabs.org/ppp-sparql/#What%20is%20the%20population%20of%20Po... [2] - https://phabricator.wikimedia.org/T68108
On Sun, Mar 6, 2016 at 11:46 AM, Sylvain Zimmer sylvain@sylvainzimmer.com wrote:
Hi,
Some of you may be familiar with http://commoncrawl.org ; they are doing an excellent job of making large crawls of the web accessible to everyone.
I've been working on an open search engine based on these crawls for a while, and I would love to have your feedbacks on the project: https://about.commonsearch.org/
Specifically, I would be curious to know what you would consider to be the best possible integration of Wikipedia & Wikidata in a general search engine?
As a first step, we have just started using the "official website" property from Wikidata and we are considering importing the Wikipedia abstracts next (https://github.com/commonsearch/cosr-back/issues/11).
I'm looking forward to your feedbacks... or contributions! :-)
Thanks in advance,
PS: A few wikimedians recommended me to post on wikitech-l to keep the focus on the technical aspects of the project and hopefully avoid linking this project in any way to the KE stuff, which it actually predates by far (https://news.ycombinator.com/item?id=6209088).
-- Sylvain Zimmer http://sylvinus.org
Wikitech-l mailing list Wikitech-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wikitech-l
wikitech-l@lists.wikimedia.org