Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Data Driven User Emulator

This software generates realistic network traffic to test intrusion detection systems. The most realistic traffic is generated when software is used to drive actual applications, thereby behaving like a real user. Tools that generate such realistic user behavior are called user emulators. However, no existing user emulators use models based on data science and real user data, and they suffer from a decrease in the fidelity of the generated traffic. We develop a user emulator that uses real user data and data science to generate higher fidelity emulation and increase the accuracy of our experimental results.

Oesch, TimothyS.↗

Sandia Academic Alliance Program Collaboration Report: 2020-2021 Accomplishments

University partnerships play an essential role in sustaining Sandia’s vitality as a national laboratory. The SAA is an element of Sandia’s broader University Partnerships program, which facilitates recruiting and research collaborations with dozens of universities annually. The SAA program has two three-year goals. SAA aims to realize a step increase in hiring results, by growing the total annual inexperienced hires from each out-of-state SAA university. SAA also strives to establish and sustain strategic research partnerships by establishing several federally sponsored collaborations and multi-institutional consortiums in science & technology (S&T) priorities such as autonomy, advanced computing, hypersonics, quantum information science, and data science. The SAA program facilitates access to talent, ideas, and Research & Development facilities through strong university partnerships. Earlier this year, the SAA program and campus executives hosted John Myers, Sandia’s former Senior Director of Human Resources (HR) and Communications, and senior-level staff at Georgia Tech, U of Illinois, Purdue, UNM, and UT Austin. These campus visits provided an opportunity to share the history of the partnerships from the university leadership, tours of research facilities, and discussions of ongoing technical work and potential recruiting opportunities. These visits also provided valuable feedback to HR management that will help Sandia realize a step increase in hiring from SAA schools. The 2020-2021 Collaboration Report is a compilation of accomplishments in 2020 and 2021 from SAA and Sandia’s valued SAA university partners.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

DevOps and Data: Faster-Time-to-Knowledge through SageOps, MLOps, and DataOps

This report describes the approach, investigation, and prototyping efforts to develop an efficient, reusable methodology and reference framework for applying DevOps to disparate data for data science and data analytics at scale, based on focused application of this methodology and reusable reference framework within Sandia National Laboratories’ Pulsed Power community. Additionally, this report reviews: engineered instantiation of the reference framework used for development and production solutions, our experiences and results in using the reference framework, and future plans regarding research and development.

97 MATHEMATICS AND COMPUTING↗

An Indicator-based Approach to Sustainable Management of Natural Resources (Chapter 12)

Assessing the sustainability of natural resource management choices for agricultural and forest lands requires quantification of potential changes to a set of environmental and socioeconomic indicators selected to characterize reference scenarios relative to projected future scenarios. Correctly framing the questions with local stakeholders is a critical first step in the sustainability assessment, and the questions that can be addressed are often limited by data availability. Selecting and prioritizing indicators with stakeholders to address their needs and concerns improves the likelihood of investment in monitoring and evaluation of those indicators over time. Computational techniques for analyzing interactions between the selected indicators are inherently affected by the scales and formats of the assembled indicator datasets. Data analytics have the potential to improve understanding of the potential synergies and tradeoffs involved with meeting multiple environmental and socioeconomic goals simultaneously, but timely and appropriate indicator datasets are not always available—even in this new era of “big data.” Continued improvements in data science and data analytics are needed to broaden understanding and acceptance of problems and to provide valuable information for natural resource management. Advances in these areas will enable society to design future landscapes that meet multiple objectives, including the provisioning of agricultural and forest resources along with a variety of ecosystem services (e.g., clean water and healthy soils).

Parish, Esther↗

Tutorial: Lessons Learned for Behavior Analysts from Data Scientists

Big data is a computing term used to refer to large and complex data sets, typically consisting of terabytes or more of diverse data that is produced rapidly. The analysis of such complex data sets requires advanced analysis techniques with the capacity to identify patterns and abstract meanings from the vast data. The field of data science combines computer science with mathematics/statistics and leverages artificial intelligence, in particular machine learning, to analyze big data. This field holds great promise for behavior analysis, where both clinical and research studies produce large volumes of diverse data at a rapid pace (i.e., big data). This article presents basic lessons for the behavior analytic researchers and clinicians regarding integration of data science into the field of behavior analysis. We provide guidance on how to collect, protect, and process the data, while highlighting the importance of collaborating with data scientists to select a proper machine learning model that aligns with the project goals and develop models with input from human experts. Here, we hope this serves as a guide to support the behavior analysts interested in the field of data science to advance their practice or research, and helps them avoid some common pitfalls.

42 ENGINEERING↗

2022 American Conference on Neutron Scattering (ACNS 2022)

The 11th American Conference on Neutron Scattering (ACNS 2022) will be held on June 5-9, 2022, in Boulder, CO. The Conference will provide essential information on the breadth and depth of current neutron-related research worldwide. Hosted by the Neutron Scattering Society of America, this year’s Conference will feature a combination of invited and contributed talks, poster sessions, and tutorials. Topics of the conference are: Advances in Neutron Facilities, Instrumentation and Software: Developments in sources, instrumentation, sample environments and control software. Hard Condensed Matter: Magnetism, correlated metals, quantum/topological materials, superconductors, ferroelectrics, multiferroics, glasses, and disorder phenomena. Submissions outlining examples of neutron scattering in industrial and engineering applications involving hard condensed matter systems are also encouraged. Soft Matter: Neutron studies of soft materials and related fields including in situ and in operando studies. Polymers, surfactants, emulsions, gels, nanoparticles, colloidal suspensions and more. Submissions of computational studies or applications of machine learning beneficial to neutron scattering experiments, as well as examples of neutron scattering in industrial and engineering applications are strongly encouraged. Biology, Biophysics and Biotechnology: Neutron studies of biological and biologically relevant systems. Proteins, bio membranes, biological assemblies, natural materials, nucleic acids, drug-delivery platforms and biomedical systems. Submissions of computational studies or applications of machine learning beneficial to biological neutron scattering experiments, as well as examples of neutron scattering in applied research involving biological systems, are strongly encouraged. Materials Chemistry and Energy: Neutron-based studies of functional materials and materials for energy applications. Examples include porous materials such as metal organic frameworks (MOFs), zeolites; phosphors; novel pigments; electrolytes; catalysts; ionic conductors/cathode materials; photovoltaic materials (hybrid perovskites); thermoelectrics; magnetocalorics/electrocalorics. Structural Materials and Engineering: Neutron scattering studies of materials and engineering processes including structural materials, concrete and metals, as well as engineering processes including combustion, corrosion, additive manufacturing, and others. Neutron Physics: Fundamental physical studies of the neutron and related areas. Emerging Applications in Neutron Scattering: Machine Learning and Data Science: Advances in computing power have contributed to rapidly evolving machine learning and data science fields that can be leveraged to the benefit of the neutron scattering community. The purpose of this session is to highlight recent advances in machine learning and data science and to serve as the foundation of a parallel data and computation track highlighting computation advances and applications in neutron scattering throughout the conference.

36 MATERIALS SCIENCE↗

47 Tuc in Rubin Data Preview 1. Exploring Early LSST Data and Science Potential

We present analyses of the early data from Rubin Observatory’s Data Preview 1 (DP1) for the field of the globular cluster 47 Tuc. The DP1 data set for 47 Tuc includes four nights of observations from the Rubin Commissioning Camera (LSSTComCam), covering multiple bands (ugriy). We address challenges of crowding in the inner region of the cluster and toward the SMC in DP1, and demonstrate improved star–galaxy separation by fitting fifth-degree polynomials to the stellar loci in color–color diagrams and applying multidimensional sigma clipping. We compile a catalog of 3576 probable 47 Tuc member stars selected via a combination of isochrone, Gaia proper-motion, and color–color space matched filtering. We explore the sources of photometric scatter in the 47 Tuc color–color sequence, evaluating contributions from various potential sources, including differential extinction within the cluster. Finally, of the 72 well-characterized variables in the field, we recover three known variable stars, including two RR Lyrae and one eclipsing binary, in the coadd-based object catalog, and identify 62 in the difference image-based object catalog. Although the DP1 lightcurves have sparse temporal sampling, they appear to follow the patterns of densely sampled literature lightcurves well. Despite some data limitations for crowded-field stellar analysis, DP1 demonstrates the promising scientific potential for future LSST data releases.

Choi, Yumi [NSF National Optical-Infrared Astronom↗

A proximal trust-region method for nonsmooth optimization with inexact function and gradient evaluations

Many applications require minimizing the sum of smooth and nonsmooth functions. For example, basis pursuit denoising problems in data science require minimizing a measure of data misfit plus an $\ell^1$-regularizer. Similar problems arise in the optimal control of partial differential equations (PDEs) when sparsity of the control is desired. Here, we develop a novel trust-region method to minimize the sum of a smooth nonconvex function and a nonsmooth convex function. Our method is unique in that it permits and systematically controls the use of inexact objective function and derivative evaluations. When using a quadratic Taylor model for the trust-region subproblem, our algorithm is an inexact, matrix-free proximal Newton-type method that permits indefinite Hessians. We prove global convergence of our method in Hilbert space and demonstrate its efficacy on three examples from data science and PDE-constrained optimization.

97 MATHEMATICS AND COMPUTING↗

Science-Driven Data Management for Multi-Tiered Storage (Final Report)

Scientific discovery at the exascale will not be possible without significant new research in the management, storage and retrieval over the long lifespan of the extreme amounts of data that will be produced. Our thesis is that adding application level knowledge about data to guide the actions of the storage system provides substantial benefits to the organization, storage, and access to extreme scale data, resulting in improved productivity for computational science. In this project we will demonstrate novel techniques to facilitate efficient mapping of data objects, even partitioning individual variables, from the user space onto multiple storage tiers, and enable application-guided data reductions and transformations to address capacity and bandwidth bottlenecks. Our goal is to address the associated Input/ Output (I/O) and storage challenges in the context of current and emerging storage landscapes, and expedite insights into mission critical scientific processes.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

LDM-151: Data Management Science Pipelines Design

The LSST Science Requirements Document (the LSST SRD) specifies a set of data product guidelines, designed to support science goals envisioned to be enabled by the LSST observing program. Following these guidelines, the details of these data products have been described in the LSST Data Products Definition Document (DPDD), and captured in a formal flow-down from the SRD via the LSST System Requirements (LSR), Observatory System Specifications (OSS), to the Data Management System Requirements (DMSR). The LSST Data Management subsystem's responsibilities include the design, implementation, deployment and execution of software pipelines necessary to generate these data products. This document describes the design of the scientific aspects of those pipelines.

79 ASTRONOMY AND ASTROPHYSICS↗

DSI Python API Demo May 2023 [Slides]

Data Science Infrastructure (DSI) is developing searchable databases and workflows for simulation and experimental data derived from ASC clients. Short term goal: Make data easily accessible through metadata indexing and querying while respecting data permissions. Longer term goal: Use this data for data science activities.

97 MATHEMATICS AND COMPUTING↗

Bioinformatic teaching resources - for educators, by educators - using KBase, a free, user-friendly, open source platform

Over the past year, biology educators and staff at the Department of Energy Systems Biology Knowledgebase (KBase) initiated a collaborative effort to develop a curriculum for bioinformatics education. KBase is a free and easily accessible data science platform that integrates many bioinformatics resources into a graphical user interface built upon reproducible analysis notebooks. KBase held conversations with college and high school instructors to understand how KBase could potentially support their educational goals. These conversations morphed into a working group of biological and data science instructors that adapted the KBase platform to their curriculum needs, specifically around concepts in Genomics, Metagenomics, Pangenomics, and Phylogenetics. The KBase Educators Working Group developed modular, adaptable, and customizable instructional units. Each instructional module contains teaching resources, publicly available data, analysis tools, and markdown capability to tailor instructions and learning goals for each class. The online user interface enables students to conduct hands-on data science research and analyses without requiring programming skills or their own computational resources (these are provided by KBase). Alongside these resources, KBase continues to work with instructors, supporting the development of additional curriculum modules. For anyone new to the platform, KBase, and the growing KBase Educators Organization, provides a community network, accompanied by community-sourced guidelines, instructional templates, and peer support to use KBase within a classroom whether virtual or in-person.

59 BASIC BIOLOGICAL SCIENCES↗

Sample Identifiers and Metadata to Support Data Management and Reuse in Multidisciplinary Ecosystem Sciences

Physical samples are foundational entities for research across biological, Earth, and environmental sciences. Data generated from sample-based analyses are not only the basis of individual studies, but can also be integrated with other data to answer new and broader-scale questions. Ecosystem studies increasingly rely on multidisciplinary team-science to study climate and environmental changes. While there are widely adopted conventions within certain domains to describe sample data, these have gaps when applied in a multidisciplinary context. In this study, we reviewed existing practices for identifying, characterizing, and linking related environmental samples. We then tested practicalities of assigning persistent identifiers to samples, with standardized metadata, in a pilot field test involving eight United States Department of Energy projects. Participants collected a variety of sample types, with analyses conducted across multiple facilities. We address terminology gaps for multidisciplinary research and make recommendations for assigning identifiers and metadata that supports sample tracking, integration, and reuse. Furthermore, our goal is to provide a practical approach to sample management, geared towards ecosystem scientists who contribute and reuse sample data.

54 ENVIRONMENTAL SCIENCES↗

Developing Fluorescence-Based Sensors to Support Rare Earth Element Separation

Rare earth elements (REEs) are essential to most renewable energy technologies. Unfortunately, as we transition to sustainable energy production, the demand for REEs is rapidly growing well beyond current rates of production. As a result, novel means of efficient, scalable, and easily adaptable methods for processing primary and recycle feedstocks are needed. Development and integration of sensors for highly selective in-line monitoring can support more efficient design and testing of such novel separation processes, as well as more cost-effective deployment of those separation flowsheets. Work here will explore the application of fluorescence spectroscopy, a highly sensitive and selective technique, to quantify multiple lanthanides in complex mixtures including known interferents or quenching agents. Results include identification of the optimal excitation wavelength and the limit of detection of various rare earth elements as well as the performance of data-science-based quantification approaches in streams where “unknowns” are present. Overall, the data science tools in conjunction with optical sensor data were able to quantify analytes in the presence of other lanthanides which can be anticipated in the actual industrial stream. Here we include characterization of lanthanides in a microfluidic device similar to those used in new process development. This study demonstrates the capability of utilizing fluorescence spectroscopy to quantify analytes in a complicated solution matrix, suggesting this is a successful approach for in-line monitoring to optimize the separation efficiency in an industrial stream.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ChemML : A machine learning and informatics program package for the analysis, mining, and modeling of chemical and materials data

ChemML is an open machine learning (ML) and informatics program suite that is designed to support and advance the data-driven research paradigm that is currently emerging in the chemical and materials domain. ChemML allows its users to perform various data science tasks and execute ML workflows that are adapted specifically for the chemical and materials context. Key features are automation, general-purpose utility, versatility, and user-friendliness in order to make the application of modern data science a viable and widely accessible proposition in the broader chemistry and materials community. Finally, ChemML is also designed to facilitate methodological innovation, and it is one of the cornerstones of the software ecosystem for data-driven in silico research.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING↗

BLDAP Intro to Python/Data Science Curriculum v1

The Github repository contains the Jupyter notebooks for the intro to Python / Data Science course for Berkeley Lab Director's Apprenticeship Program (BLDAP). This course is designed for students with little to no experience in coding to learn skills in Python necessary for data science. Students utilize Jupyter notebooks throughout the course. The overall goal is for students to learn how to use Python to clean, analyze, and visualize large data sets in order to communicate effectively their conclusions about the data set. Students apply the skills they learned on actual data sets provided by researchers in Berkeley Lab.

Hales, Laurel [Lawrence Berkeley National Laborato↗

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)↗