Subject: DSCI
Term: Spring 2021
DSCI 332: Spatial Statistics for Near Surface, Surface, and Subsurface Modeling (800/9976)
TR 10:00-11:15 AM Feb 01-May 07
Yarus, J
This course is on spatial modeling of near surface, surface, and subsurface data, also known as geostatistical modeling. Spatial modeling has its origins in predictive modeling of minerals in subsurface formations, from which many examples are used in this class. Students will learn the basics of spatial models in order to understand how they are built from various data types and how their uncertainties are assessed and risk reduced. Students will be expected to learn the rudimentary navigation of R Studio, execute pre-written publically available R code (provided), and make simple modifications. Graduate students will be expected to learn the above and develop a 10 week modeling project focused on the use of spatial modeling methods with R using data relevant to their specific discipline or interest. These projects will include preparing datasets to be executed in R code scripts. Resulting scripts will be placed in a git repository for use by other students as open source resources along with documentation demonstrating the reproducible spatial modeling science and analyses for these problems.
Geostatistical (spatial) mapping is applicable across many disciplines. Examples of graduate projects from previous classes include subsurface modeling (geology), earthquake mapping (geophysics/civil engineering), soil stability modeling (civil engineering), aquifer characterization (hydrology), and pollution/contaminant mapping (environmental studies/medicine).
Offered as DSCI 332 and DSCI 432.
DSCI 352: Applied Data Science Research (800/4652)
F 12:45-01:45 PM Feb 01-May 07
Bruckman, L
This is a project based data science research class, in which project teams identify a research project under the guidance of a domain expert professor. The research is structured as a data analysis project including the 6 steps of developing a reproducible data science project, including 1: Define the ADS question, 2: Identify, locate, and/or generate the data 3: Exploratory data analysis 4: Statistical modeling and prediction 5: Synthesizing the results in the domain context 6: Creation of reproducible research, Including code, datasets, documentation and reports. During the course special topic lectures will include Ethics, Privacy, Openness, Security, Ethics. Value. The M section of DSCI 352 is for students focusing on Materials Data Science.
Offered as DSCI 352, DSCI 352M and DSCI 452.
DSCI 352M: Applied Data Science Research (800/4802)
F 12:45-01:45 PM Feb 01-May 07
Bruckman, L
This is a project based data science research class, in which project teams identify a research project under the guidance of a domain expert professor. The research is structured as a data analysis project including the 6 steps of developing a reproducible data science project, including 1: Define the ADS question, 2: Identify, locate, and/or generate the data 3: Exploratory data analysis 4: Statistical modeling and prediction 5: Synthesizing the results in the domain context 6: Creation of reproducible research, Including code, datasets, documentation and reports. During the course special topic lectures will include Ethics, Privacy, Openness, Security, Ethics. Value. The M section of DSCI 352 is for students focusing on Materials Data Science.
Offered as DSCI 352, DSCI 352M and DSCI 452.
DSCI 353: Data Science: Statistical Learning, Modeling and Prediction (800/4651)
TR 11:30-12:45 PM Feb 01-May 07
French, R
In this course, we will use an open data science tool chain to develop reproducible data analyses useful for inference, modeling and prediction of the behavior of complex systems. In addition to the standard data cleaning, assembly and exploratory data analysis steps essential to all data analyses, we will identify statistically significant relationships from datasets derived from population samples, and infer the reliability of these findings. We will use regression methods to model a number of both real-world and lab-based systems producing predictive models applicable in comparable populations. We will assemble and explore real-world datasets, use pair-wise plots to explore correlations, perform clustering, self-similarity, and logistic regression develop both fixed-effect and mixed-effect predictive models. We will introduce machine-learning approaches for classification and tree-based methods. Results will be interpreted, visualized and discussed. We will introduce the basic elements of data science and analytics using R Project open source software. R is an open-source software project with broad abilities to access machine-readable open-data resources, data cleaning and assembly functions, and a rich selection of statistical packages, used for data analytics, model development, prediction, inference and clustering. With this background, it becomes possible to start performing variable transformations for linear regression fitting and developing structural equation models, fixed-effects and mixed-effects models along with other statistical learning techniques, while exploring for statistically significant relationships. The class will be structured to have a balance of theory and practice. We'll split class into Foundation and Practicum a) Foundation: lectures, presentations, discussion b) Practicum: coding, demonstrations and hands-on data science work. The M section of DSCI 353 is for students focusing on Materials Data Science.
Offered as DSCI 353, DSCI 353M and DSCI 453.
DSCI 353M: Data Science: Statistical Learning, Modeling and Prediction (800/4749)
TR 11:30-12:45 PM Feb 01-May 07
French, R
In this course, we will use an open data science tool chain to develop reproducible data analyses useful for inference, modeling and prediction of the behavior of complex systems. In addition to the standard data cleaning, assembly and exploratory data analysis steps essential to all data analyses, we will identify statistically significant relationships from datasets derived from population samples, and infer the reliability of these findings. We will use regression methods to model a number of both real-world and lab-based systems producing predictive models applicable in comparable populations. We will assemble and explore real-world datasets, use pair-wise plots to explore correlations, perform clustering, self-similarity, and logistic regression develop both fixed-effect and mixed-effect predictive models. We will introduce machine-learning approaches for classification and tree-based methods. Results will be interpreted, visualized and discussed. We will introduce the basic elements of data science and analytics using R Project open source software. R is an open-source software project with broad abilities to access machine-readable open-data resources, data cleaning and assembly functions, and a rich selection of statistical packages, used for data analytics, model development, prediction, inference and clustering. With this background, it becomes possible to start performing variable transformations for linear regression fitting and developing structural equation models, fixed-effects and mixed-effects models along with other statistical learning techniques, while exploring for statistically significant relationships. The class will be structured to have a balance of theory and practice. We'll split class into Foundation and Practicum a) Foundation: lectures, presentations, discussion b) Practicum: coding, demonstrations and hands-on data science work. The M section of DSCI 353 is for students focusing on Materials Data Science.
Offered as DSCI 353, DSCI 353M and DSCI 453.
DSCI 432: Spatial Statistics for Near Surface, Surface, and Subsurface Modeling (800/9969)
TR 10:00-11:15 AM Feb 01-May 07
Yarus, J
This course is on spatial modeling of near surface, surface, and subsurface data, also known as geostatistical modeling. Spatial modeling has its origins in predictive modeling of minerals in subsurface formations, from which many examples are used in this class. Students will learn the basics of spatial models in order to understand how they are built from various data types and how their uncertainties are assessed and risk reduced. Students will be expected to learn the rudimentary navigation of R Studio, execute pre-written publically available R code (provided), and make simple modifications. Graduate students will be expected to learn the above and develop a 10 week modeling project focused on the use of spatial modeling methods with R using data relevant to their specific discipline or interest. These projects will include preparing datasets to be executed in R code scripts. Resulting scripts will be placed in a git repository for use by other students as open source resources along with documentation demonstrating the reproducible spatial modeling science and analyses for these problems.
Geostatistical (spatial) mapping is applicable across many disciplines. Examples of graduate projects from previous classes include subsurface modeling (geology), earthquake mapping (geophysics/civil engineering), soil stability modeling (civil engineering), aquifer characterization (hydrology), and pollution/contaminant mapping (environmental studies/medicine).
Offered as DSCI 332 and DSCI 432.
DSCI 452: Applied Data Science Research (800/4748)
F 12:45-01:45 PM Feb 01-May 07
Bruckman, L
This is a project based data science research class, in which project teams identify a research project under the guidance of a domain expert professor. The research is structured as a data analysis project including the 6 steps of developing a reproducible data science project, including 1: Define the ADS question, 2: Identify, locate, and/or generate the data 3: Exploratory data analysis 4: Statistical modeling and prediction 5: Synthesizing the results in the domain context 6: Creation of reproducible research, Including code, datasets, documentation and reports. During the course special topic lectures will include Ethics, Privacy, Openness, Security, Ethics. Value. The M section of DSCI 352 is for students focusing on Materials Data Science.
Offered as DSCI 352, DSCI 352M and DSCI 452.
DSCI 452: Applied Data Science Research (500/9971)
Bruckman, L
This is a project based data science research class, in which project teams identify a research project under the guidance of a domain expert professor. The research is structured as a data analysis project including the 6 steps of developing a reproducible data science project, including 1: Define the ADS question, 2: Identify, locate, and/or generate the data 3: Exploratory data analysis 4: Statistical modeling and prediction 5: Synthesizing the results in the domain context 6: Creation of reproducible research, Including code, datasets, documentation and reports. During the course special topic lectures will include Ethics, Privacy, Openness, Security, Ethics. Value. The M section of DSCI 352 is for students focusing on Materials Data Science.
Offered as DSCI 352, DSCI 352M and DSCI 452.
DSCI 453: Data Science: Statistical Learning, Modeling and Prediction (800/4655)
TR 11:30-12:45 PM Feb 01-May 07
French, R
In this course, we will use an open data science tool chain to develop reproducible data analyses useful for inference, modeling and prediction of the behavior of complex systems. In addition to the standard data cleaning, assembly and exploratory data analysis steps essential to all data analyses, we will identify statistically significant relationships from datasets derived from population samples, and infer the reliability of these findings. We will use regression methods to model a number of both real-world and lab-based systems producing predictive models applicable in comparable populations. We will assemble and explore real-world datasets, use pair-wise plots to explore correlations, perform clustering, self-similarity, and logistic regression develop both fixed-effect and mixed-effect predictive models. We will introduce machine-learning approaches for classification and tree-based methods. Results will be interpreted, visualized and discussed. We will introduce the basic elements of data science and analytics using R Project open source software. R is an open-source software project with broad abilities to access machine-readable open-data resources, data cleaning and assembly functions, and a rich selection of statistical packages, used for data analytics, model development, prediction, inference and clustering. With this background, it becomes possible to start performing variable transformations for linear regression fitting and developing structural equation models, fixed-effects and mixed-effects models along with other statistical learning techniques, while exploring for statistically significant relationships. The class will be structured to have a balance of theory and practice. We'll split class into Foundation and Practicum a) Foundation: lectures, presentations, discussion b) Practicum: coding, demonstrations and hands-on data science work. The M section of DSCI 353 is for students focusing on Materials Data Science.
Offered as DSCI 353, DSCI 353M and DSCI 453.
DSCI 453: Data Science: Statistical Learning, Modeling and Prediction (500/9972)
French, R
In this course, we will use an open data science tool chain to develop reproducible data analyses useful for inference, modeling and prediction of the behavior of complex systems. In addition to the standard data cleaning, assembly and exploratory data analysis steps essential to all data analyses, we will identify statistically significant relationships from datasets derived from population samples, and infer the reliability of these findings. We will use regression methods to model a number of both real-world and lab-based systems producing predictive models applicable in comparable populations. We will assemble and explore real-world datasets, use pair-wise plots to explore correlations, perform clustering, self-similarity, and logistic regression develop both fixed-effect and mixed-effect predictive models. We will introduce machine-learning approaches for classification and tree-based methods. Results will be interpreted, visualized and discussed. We will introduce the basic elements of data science and analytics using R Project open source software. R is an open-source software project with broad abilities to access machine-readable open-data resources, data cleaning and assembly functions, and a rich selection of statistical packages, used for data analytics, model development, prediction, inference and clustering. With this background, it becomes possible to start performing variable transformations for linear regression fitting and developing structural equation models, fixed-effects and mixed-effects models along with other statistical learning techniques, while exploring for statistically significant relationships. The class will be structured to have a balance of theory and practice. We'll split class into Foundation and Practicum a) Foundation: lectures, presentations, discussion b) Practicum: coding, demonstrations and hands-on data science work. The M section of DSCI 353 is for students focusing on Materials Data Science.
Offered as DSCI 353, DSCI 353M and DSCI 453.