• Home
  • Blog
  • People
  • Projects
  • Publications
  • Seminars
  • DSR Expo
  • Courses
logo1

Data Science Research

Menu
  • Home
  • Blog
  • People
  • Projects
  • Publications
  • Seminars
  • DSR Expo
  • Courses
Home › courses › Three-Course Data Science Curriculum @ UF CISE

Three-Course Data Science Curriculum @ UF CISE

Daisy Zhe Wang December 31, 2013     No Comment     courses

Daisy Zhe Wang

Daisy Zhe Wang

In order to address the growing need from both industry and academia (e.g., medical and bio informatics, financial, law enforcement, economics, decision support, social networks) for big data analytic skills including, data management, data mining, natural language processing, machine learning and data visualization, we introduce a three-course series in the Data Science Curriculum:

  1. Introduction to Data Science (both undergrad and graduate level)
  2. Advanced Topics in Data Science (graduate level only)
  3. Projects in Data Science (both undergrad and graduate level)

Due to the inter-disciplinary nature of Data Science applications, we encourage students from CS as well as other majors with CS minor to take the first and third course in the curriculum. We encourage CS graduates to take the second course to explore and push forward the frontier of Data Science technology. We will start to offer the first course in the series “Introduction to Data Science” in Spring 2014.

The aim of the first course “Introduction to Data Science” is to bring student with basic programming and data structure background to be abreast with common tools used for Data Science application development. This course will give an introduction to the basic data science techniques including programming in SQL, Map-Reduce, R, and Python. We also cover topics including relational databases, data visualization, classification, clustering, regression and parallel computing platforms. Some of tentative topics to be covered are:

Part 0: Introduction

Part 1: Data Manipulation, at Scale

  • MapReduce, Hadoop, relationship to databases, algorithms, extensions, languages
  • Databases, SQL and the relational algebra
  • Parallel databases, parallel query processing, in-database analytics
  • Key-value stores and NoSQL; tradeoffs of SQL and NoSQL

Part 2: Statistical Analytics

  • Programming in Python and R
  • Basic Data Mining
    • Basic statistical modeling, introduction to machine learning, overfitting
    • Supervised learning: Linear and Logistic Regression, Classification
    • Unsupervised learning: Clustering, Association Rule mining

Part 3: Graph/Text Data Analysis & Communicating Results

  • Graph Analytics: PageRank, community detection, recursive queries, iterative processing
  • Text Analytics: TF/IDF, conditional random fields
  • Visualization, data products, visual data analytics

Part 4: Parallel Computing

  • Concurrency and Data Decomposition
  • Message Based Parallelism – MPI
  • Thread Based Parallelism – OpenMP

The course will be mainly project based. We encourage students to form groups to develop Data Science application to compete the the 2nd UF Data Science Exposition. In the 1st UF Data Science Exposition, we received generous sponsorship from Google And Amazon. Stay tuned!

courses

 Previous Post

GPText: Greenplum Parallel Statistical Text Analysis Framework

― November 11, 2013

Next Post 

SMART Electronic Discovery

― March 14, 2014

Related Articles

DBSim: Extensible Database Simulator for Fast Prototyping In-Database Algorithms
DrugEHRQA: A Question Answering Dataset on Structured and Unstructured Electronic Health Records For Medicine Related Queries
A Brief Overview of Weak Supervision
DRUM: End-To-End Differentiable Rule Mining On Knowledge Graphs
IDTrees Data Science Challenge: 2017

Leave a Reply Cancel reply

You must be logged in to post a comment.

Recent Posts

  • DBSim: Extensible Database Simulator for Fast Prototyping In-Database Algorithms
  • DrugEHRQA: A Question Answering Dataset on Structured and Unstructured Electronic Health Records For Medicine Related Queries
  • A Brief Overview of Weak Supervision
  • DRUM: End-To-End Differentiable Rule Mining On Knowledge Graphs
  • IDTrees Data Science Challenge: 2017

Categories

  • courses
  • ecology
  • NIST and open eval
  • publications
  • research
  • research directions
  • survey
  • Uncategorized

Archives

  • February 2023
  • October 2020
  • December 2019
  • April 2019
  • December 2018
  • August 2018
  • February 2018
  • November 2017
  • June 2017
  • May 2017
  • March 2017
  • December 2016
  • October 2016
  • April 2016
  • March 2016
  • December 2015
  • November 2015
  • October 2015
  • May 2015
  • November 2014
  • October 2014
  • July 2014
  • May 2014
  • March 2014
  • December 2013
  • November 2013
  • October 2013
  • September 2013

Recent Posts

  • DBSim: Extensible Database Simulator for Fast Prototyping In-Database Algorithms
  • DrugEHRQA: A Question Answering Dataset on Structured and Unstructured Electronic Health Records For Medicine Related Queries
  • A Brief Overview of Weak Supervision
  • DRUM: End-To-End Differentiable Rule Mining On Knowledge Graphs
  • IDTrees Data Science Challenge: 2017