Tweet The UMBC WebBase corpus is a dataset of high quality English paragraphs containing over three billion words derived from the Stanford WebBase project’s February 2007 Web crawl....
Tweet Facebook engineers Xiao Li and Maxime Boucher describe the language processing techniques used to implement Facebook’s graph search in a recent post on the Facebook Engineering page...
Tweet Barbara Starr posts at Search Engine Land about progress the notion of “semantic search” has made in the past three years, Semantic and Graph-Based Search: The Future Face Of Search:...
Tweet Heather McIlvaine from enterprise software giant SAP blogs about open data: “How are mobile apps, Big Data, and civic hacking changing the nature of open data in government? The Center for...
Tweet A post in Micrsoft’s Bing blog, Understand Your World with Bing, announced that an update to their Satori knowledge base allows Bing to do a better job of identifying queries that...
Tweet Google released the Wikilinks Corpus, a collection of 40M disambiguated mentions from 10M web pages to 3M Wikipedia pages. This data can be used to train systems that do entity linking...
Tweet Google Sets was a the result of a early Google research project that ended in 2011. The idea was to be able to recognize the similarity of a set of terms (e.g., python, lisp and...
Tweet The popular KDnuggets news site for analytics, data mining and data science asked their visitors “What will replace “Big Data” as a hot buzzword?” and the most popular choice was “smart...
Tweet Computing semantic similarity between words and phrases has important applications in natural language processing, information retrieval, and artificial intelligence. There are two...
Tweet Google has added an “entity disambiguation” feature along with auto-complete when you type in your search query. For example, when I search for George Bush, I get the following additional...