• Home
  • Blog
  • People
  • Projects
  • Publications
  • Seminars
  • DSR Expo
  • Courses
logo1

Data Science Research

Menu
  • Home
  • Blog
  • People
  • Projects
  • Publications
  • Seminars
  • DSR Expo
  • Courses

MADden: Query-Driven Statistical Text Analytics

In many domains, structured data and unstructured text are both important natural resources to fuel data analysis. Statistical text analysis needs to be performed over text data to extract structured information for further query processing. Typically, developers will need to connect multiple tools to build off-line batch processes to perform text analytic tasks. MADden is an integrated system developed for relational database systems such as PostgreSQL and Greenplum for real-time ad hoc query processing over structured and unstructured data. MADden implements four important text analytic functions that we have contributed to the MADlib open source library for textual analytics. In this demonstration, we will show the capability of the MADden text analytic library using computational journalism as the driving application. We show real-time declarative query processing over multiple data sources with both structured and text information.

Authors: 
Christan Grant, Jordan Gumbs, Kun Li, Daisy Zhe Wang, George Chitouras

Bibtex:

@inproceedings{Grant:2012:MQS:2396761.2398746,
 author = {Grant, Christan Earl and Gumbs, Joir-dan and Li, Kun and Wang, Daisy Zhe and Chitouras, George},
 title = {MADden: query-driven statistical text analytics},
 booktitle = {Proceedings of the 21st ACM international conference on Information and knowledge management},
 series = {CIKM '12},
 year = {2012},
 isbn = {978-1-4503-1156-4},
 location = {Maui, Hawaii, USA},
 pages = {2740--2742},
 numpages = {3},
 url = {http://doi.acm.org/10.1145/2396761.2398746},
 doi = {10.1145/2396761.2398746},
 acmid = {2398746},
 publisher = {ACM},
 address = {New York, NY, USA},
 keywords = {databases, query-driven, text analytics},
}

Download:
[pdf]

Recent Posts

  • DBSim: Extensible Database Simulator for Fast Prototyping In-Database Algorithms
  • DrugEHRQA: A Question Answering Dataset on Structured and Unstructured Electronic Health Records For Medicine Related Queries
  • A Brief Overview of Weak Supervision
  • DRUM: End-To-End Differentiable Rule Mining On Knowledge Graphs
  • IDTrees Data Science Challenge: 2017

Categories

  • courses
  • ecology
  • NIST and open eval
  • publications
  • research
  • research directions
  • survey
  • Uncategorized

Archives

  • February 2023
  • October 2020
  • December 2019
  • April 2019
  • December 2018
  • August 2018
  • February 2018
  • November 2017
  • June 2017
  • May 2017
  • March 2017
  • December 2016
  • October 2016
  • April 2016
  • March 2016
  • December 2015
  • November 2015
  • October 2015
  • May 2015
  • November 2014
  • October 2014
  • July 2014
  • May 2014
  • March 2014
  • December 2013
  • November 2013
  • October 2013
  • September 2013