• Home
  • Blog
  • People
  • Projects
  • Publications
  • Seminars
  • DSR Expo
  • Courses
logo1

Data Science Research

Menu
  • Home
  • Blog
  • People
  • Projects
  • Publications
  • Seminars
  • DSR Expo
  • Courses
Home › 2013
  • Streaming Fact Extraction for Wikipedia Entities at Web-Scale

    May 12, 2014     No Comment     Uncategorized

    Morteza Shahriari Nia, Christan Grant, Yang Peng Wikipedia.org (WP) is the largest and most popular general reference work on the Internet. Presently, there is considerable time lag between the publication of an event and its citation in WP. The median time lag for a sample of about 60K web pages cited by WP articles in

    Read more »

  • CASTLE: Crowd-Assisted System for Textual Labeling & Extraction

    October 20, 2013     No Comment    

    The amount of text data has been growing exponentially and with it the demand for improved information extraction (IE) efforts to analyze and query such data. While automatic IE systems have proven useful in controlled experiments, in practice the gap between machine learning extraction and human extraction is still quite large. In this paper, we propose a system that uses crowdsourcing

    Read more »

  • GPText: Greenplum Parallel Statistical Text Analysis Framework

    August 26, 2013     Comment Closed    

    Many companies keep large amounts of text data inside of relational databases. Several challenges exist in using state-of-the-art systems to perform analysis on such datasets. First, expensive big data transfer cost must be paid up front to move data between databases and analytics systems. Second, many popular text analytics packages do not scale up to production sized datasets.

    Read more »

  • Web-Scale Knowledge Inference Using Markov Logic Networks

    August 26, 2013     Comment Closed    

    In this paper, we present our on-going work on ProbKB, a PROBabilistic Knowledge Base constructed from web-scale extracted entities, facts, and rules represented as a Markov logic network (MLN). We aim at web-scale MLN inference by designing a novel relational model to represent MLNs and algorithms that apply rules in batches. Errors are handled in a principled and elegant manner to avoid error

    Read more »

  • Knowledge Extraction and Outcome Prediction using Medical Notes

    August 26, 2013     Comment Closed    

    The increasing use of electronic health records (EHR) has allowed for an unprecedented ability to perform analysis on patient data. By training a number of statistical machine learning classifiers over the unstructured text found in admission notes and operating procedures, prediction of a surgical procedure’s outcome can be performed. We extend an initial bag-of-words model to a bag-of-concepts model, which uses cTakes and UMLS

    Read more »

Recent Posts

  • DBSim: Extensible Database Simulator for Fast Prototyping In-Database Algorithms
  • DrugEHRQA: A Question Answering Dataset on Structured and Unstructured Electronic Health Records For Medicine Related Queries
  • A Brief Overview of Weak Supervision
  • DRUM: End-To-End Differentiable Rule Mining On Knowledge Graphs
  • IDTrees Data Science Challenge: 2017

Categories

  • courses
  • ecology
  • NIST and open eval
  • publications
  • research
  • research directions
  • survey
  • Uncategorized

Archives

  • February 2023
  • October 2020
  • December 2019
  • April 2019
  • December 2018
  • August 2018
  • February 2018
  • November 2017
  • June 2017
  • May 2017
  • March 2017
  • December 2016
  • October 2016
  • April 2016
  • March 2016
  • December 2015
  • November 2015
  • October 2015
  • May 2015
  • November 2014
  • October 2014
  • July 2014
  • May 2014
  • March 2014
  • December 2013
  • November 2013
  • October 2013
  • September 2013

Recent Posts

  • DBSim: Extensible Database Simulator for Fast Prototyping In-Database Algorithms
  • DrugEHRQA: A Question Answering Dataset on Structured and Unstructured Electronic Health Records For Medicine Related Queries
  • A Brief Overview of Weak Supervision
  • DRUM: End-To-End Differentiable Rule Mining On Knowledge Graphs
  • IDTrees Data Science Challenge: 2017