• Home
  • Blog
  • People
  • Projects
  • Publications
  • Seminars
  • DSR Expo
  • Courses
logo1

Data Science Research

Menu
  • Home
  • Blog
  • People
  • Projects
  • Publications
  • Seminars
  • DSR Expo
  • Courses

A Machine Learning Based Topic Exploration and Categorization on Survey

This paper describes an automatic topic extraction, categorization, and relevance ranking model for multilingual surveys and questions that exploits machine learning algorithms such as topic modeling and fuzzy clustering. Automatically generated question and survey categories are used to build question banks and category-specific survey templates. First, we describe different pre-processing steps we considered for removing noise in the multilingual survey text. Second, we explain our strategy to automatically extract survey categories from surveys based on topic models. Third, we describe different methods to cluster questions under survey categories and group them based on relevance. Last, we describe our experimental results on a large group of unique, real-world survey datasets from the German, Spanish, French, and Portuguese languages and our refining methods
to determine meaningful and sensible categories for building question banks. We conclude this document with possible enhancements to the current system and impacts in the business domain.

Authors: 
Clint P. George, Daisy Zhe Wang, Joseph N. Wilson, Liana M. Epstein, Philip Garland, Annabell Suh

Bibtex:

@inproceedings{George:2012:MLB:2456693.2453720,
 author = {George, Clint P. and Wang, Daisy Zhe and Wilson, Joseph N. and Epstein, Liana M. and Garland, Philip and Suh, Annabell},
 title = {A Machine Learning Based Topic Exploration and Categorization on Surveys},
 booktitle = {Proceedings of the 2012 11th International Conference on Machine Learning and Applications - Volume 02},
 series = {ICMLA '12},
 year = {2012},
 isbn = {978-0-7695-4913-2},
 pages = {7--12},
 numpages = {6},
 url = {http://dx.doi.org/10.1109/ICMLA.2012.132},
 doi = {10.1109/ICMLA.2012.132},
 acmid = {2453720},
 publisher = {IEEE Computer Society},
 address = {Washington, DC, USA},
 keywords = {topic modeling, survey clustering, fuzzy clustering, categorization},
}

Download:
[pdf]

Recent Posts

  • DBSim: Extensible Database Simulator for Fast Prototyping In-Database Algorithms
  • DrugEHRQA: A Question Answering Dataset on Structured and Unstructured Electronic Health Records For Medicine Related Queries
  • A Brief Overview of Weak Supervision
  • DRUM: End-To-End Differentiable Rule Mining On Knowledge Graphs
  • IDTrees Data Science Challenge: 2017

Categories

  • courses
  • ecology
  • NIST and open eval
  • publications
  • research
  • research directions
  • survey
  • Uncategorized

Archives

  • February 2023
  • October 2020
  • December 2019
  • April 2019
  • December 2018
  • August 2018
  • February 2018
  • November 2017
  • June 2017
  • May 2017
  • March 2017
  • December 2016
  • October 2016
  • April 2016
  • March 2016
  • December 2015
  • November 2015
  • October 2015
  • May 2015
  • November 2014
  • October 2014
  • July 2014
  • May 2014
  • March 2014
  • December 2013
  • November 2013
  • October 2013
  • September 2013