Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

NSDS Nuclear Science Data Solutions [Poster]

NSDS is the adept coordination of Los Alamos National Laboratory's nuclear science data, including the collection, integration, storage and access. The poster highlights historical milestones.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Revolutionizing Energetic Materials Discovery and Design: The Role of Data Science and Machine Learning

Here this Special Issue of Propellants, Explosives, Pyrotechnics (PEP) is focused on energetic materials discovery and design using Data Science and Machine Learning (DS&ML). The application of DS&ML has proven to be transformative in many areas, where it has been shown to expedite analysis, enable extraction of greater quantities of information from datasets, and guide experiments. However, energetic materials and their applications present unique challenges that often hinder the use of standardized tools and practices. In spite of these challenges, important and compelling advancements are being made toward data-directed research in energetics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Square Kilometre Array Science Data Challenge 1: analysis and results

ABSTRACT As the largest radio telescope in the world, the Square Kilometre Array (SKA) will lead the next generation of radio astronomy. The feats of engineering required to construct the telescope array will be matched only by the techniques developed to exploit the rich scientific value of the data. To drive forward the development of efficient and accurate analysis methods, we are designing a series of data challenges that will provide the scientific community with high-quality data sets for testing and evaluating new techniques. In this paper, we present a description and results from the first such Science Data Challenge 1 (SDC1). Based on SKA MID continuum simulated observations and covering three frequencies (560, 1400, and 9200 MHz) at three depths (8, 100, and 1000 h), SDC1 asked participants to apply source detection, characterization, and classification methods to simulated data. The challenge opened in 2018 November, with nine teams submitting results by the deadline of 2019 April. In this work, we analyse the results for eight of those teams, showcasing the variety of approaches that can be successfully used to find, characterize, and classify sources in a deep, crowded field. The results also demonstrate the importance of building domain knowledge and expertise on this kind of analysis to obtain the best performance. As high-resolution observations begin revealing the true complexity of the sky, one of the outstanding challenges emerging from this analysis is the ability to deal with highly resolved and complex sources as effectively as the unresolved source population.

Bonaldi, A.↗

Application of Data Science and Engineering

Metal additive manufacturing (AM) processes exhibit significant variability in the quality and properties of components that are produced. This variability has prevented the widespread adoption of AM in industry. The need for more advanced and descriptive process monitoring, part qualification, and process control has led to an increasing number of sensors on machines and subsequent data to analyze. Increasingly, data science principles are being leveraged in each of these domains in order to process this data and better understand the causes of variability and the corresponding quality inconsistencies that occur in additive manufacturing.

Halsey, William↗

In the Mix : A Workshop Merging Computational Chemistry and Electrochemistry Alongside Data Science

As chemistry expands to more complex and interdisciplinary areas, a new generation of diverse researchers must engage with science and learn effective cross-disciplinary collaboration and communication. To these ends, we designed and implemented In the Mix, a graduate student-led, two-day workshop for undergraduate students promoting collaborative science in the context of energy storage innovations. Here, the interactive workshop was designed for future and emerging researchers to gain hands-on experience with data science, computational chemistry, and electrochemistry techniques that are critical for developing materials for battery technologies. Participants also visited commercial renewable energy facilities to help them connect discovery-based research with industry and broader societal considerations. The workshop content and structure ensured that participants experienced the interrelatedness of the fields and understood the importance of collaborative research to yield scientific advances with real-world applications. An external team evaluated the workshop and participants’ perceptions of their experiences. While our research context was energy storage, the workshop goals and outcomes are applicable to other contexts. Interdisciplinary, experiential workshops are a key avenue to broadening participation in science and research, and the ideas presented here can be readily modified for other scientific contexts and/or incorporated as broader impact activities.

25 ENERGY STORAGE↗

Data Science Enabled Enabled Discovery of Superconductors (Final Progress Report)

This Final Technical Report describes efforts by 4 PIs at the University of Florida (Peter Hirschfeld, Richard Hennig, Greg Stewart and James Hamlin), over the period September 2019-August 2023, to use data science and machine learning techniques to discover new conventional superconductors. The PIs constructed a discovery loop with two theorists and two experimentalists to: develop algorithms to machine learn descriptors correlating strongly with the critical temperature Tc (PI's Peter Hirschfeld, UF Physics and Richard Hennig, UF Materials Science and En), synthesize and measure properties of promising materials, and feed back the knowledge gained into the prediction algorithm. This work was motivated by the theoretical prediction and experimental discovery of high-pressure, high-pressure hydride superconductors, and to find ways to recreate the high critical temperatures in these systems at ambient pressure. Highlights from the grant include: 1) a new equation for Tc in terms of moments of the electron-phonon spectral function, improving on the so-called Allen-Dynes equation (1975); 2) study of the metastable A15 superconductor Nb3Si, formed under explosive compression at ~1000GPa to determine the kinetic barrier to the ground state structure; 3) the development of ultra-fast machine-learned atomic potentials for molecular dynamics, and 4) the discovery of superconductivity at 19K in WB2 arising from metastable defect structures in the crystal.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

HL-LHC Computing Review Stage 2, Common Software Projects: Data Science Tools for Analysis

This paper was prepared by the HEP Software Foundation (HSF) PyHEP Working Group as input to the second phase of the LHCC review of High-Luminosity LHC (HL-LHC) computing, which took place in November, 2021. It describes the adoption of Python and data science tools in HEP, discusses the likelihood of future scenarios, and recommendations for action by the HEP community.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Using Data Science Tools to Reveal and Understand Subtle Relationships of Inhibitor Structure in Frontal Ring-Opening Metathesis Polymerization

The rate of frontal ring-opening metathesis polymerization (FROMP) using the Grubbs generation II catalyst is impacted by both the concentration and choice of monomers and inhibitors, usually organophosphorus derivatives. Herein we report a data-science-driven workflow to evaluate how these factors impact both the rate of FROMP and how long the formulation of the mixture is stable (pot life). Using this workflow, we built a classification model using a single-node decision tree to determine how a simple phosphine structural descriptor (V bur-near ) can bin long versus short pot life. Additionally, we applied a nonlinear kernel ridge regression model to predict how the inhibitor and selection/concentration of comonomers impact the FROMP rate. Furthermore, the analysis provides selection criteria for material network structures that span from highly cross-linked thermosets to non-cross-linked thermoplastics as well as degradable and nondegradable materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A General Materials Data Science Framework for Quantitative 2D Analysis of Particle Growth from Image Sequences

Abstract Phase transformations are a challenging problem in materials science, which lead to changes in properties and may impact performance of material systems in various applications. We introduce a general framework for the analysis of particle growth kinetics by utilizing concepts from machine learning and graph theory. As a model system, we use image sequences of atomic force microscopy showing the crystallization of an amorphous fluoroelastomer film. To identify crystalline particles in an amorphous matrix and track the temporal evolution of the particle dispersion, we have developed quantitative methods of 2D analysis. 700 image sequences were analyzed using a neural network architecture, achieving 0.97 pixel-wise classification accuracy as a measure of the correctly classified pixels. The growth kinetics of isolated and impinged particles were tracked throughout time using these image sequences. The relationship between image sequences and spatiotemporal graph representations was explored to identify the proximity of crystallites from each other. The framework enables the analysis of all image sequences without the requirement of sampling for specific particles or timesteps for various materials systems.

36 MATERIALS SCIENCE↗

Advanced Data Science Model for Detecting Intelligent Malware

This study focused on developing a robust artificial intelligence (AI) model capable of detecting and characterizing advanced malware in Internet of Things (IoT) devices using network data. By analyzing network traffic with various machine learning (ML) models, our AI model can identify and characterize malicious activities to significantly improve malware detection accuracy and reliability as compared to traditional methods. The developed AI/ML model was trained using network data from IoT devices, leveraging classifiers such as Random Forest, Gradient Boosting, AdaBoost, and others to optimize detection performance. This project demonstrates a scalable framework for real-time malware detection and characterization in IoT networks, capable of identifying infected devices and facilitating the necessary steps to remove or isolate them, thereby preventing further infections. Although digital twin (DT) integration is not yet implemented in the current model, it represents a promising future enhancement. By creating a virtual replica of physical IoT devices, DT technology would allow for real-time monitoring and analysis without directly accessing operational technology, thus reducing the risk of compromising or reducing the performance of actual devices. This integration would further enhance the security of IoT ecosystems, combining AI technology to better flag and detect indications of malware-infected devices within a nuclear system environment.

42 ENGINEERING↗

Data Science Techniques, Assumptions, and Challenges in Alloy Clustering and Property Prediction

Data analytics methods have been increasingly applied to understanding materials chemistry, processing due to the manufacturing approach, and uni-axial and cyclic property relationships in the highly complex space of alloy design. There are several benefits to applying data analytics to this space, including the ability to manage non-linearities in the responses of the alloy attributes and the resulting mechanical properties. However, key difficulties in applying and understanding the results of data analytics include the often lack of reported assumptions and data processing steps necessary to improve interpretation and reproducibility in derived results. In this work, the methods used to generate clustering and correlation analyses for experimental 9% Cr ferritic-martensitic steel data were investigated and the resulting implications for mechanical property predictions were assessed. This work uses principal component analysis, partitioning around medoids, t-SNE, and k-means clustering to investigate trends in composition, processing and microstructure information with creep and tensile properties, building on work done previously using a smaller version of the same dataset. The initial assumptions, preprocessing steps and methods are investigated and outlined in order to depict the fine level of detail required to convey the steps taken to process data and produce analytical results. Here, the variations in the resulting analyses are explored due to the influence of new and more varied data.

36 MATERIALS SCIENCE↗

Advanced Computing, Data Science, and Artificial Intelligence Research Opportunities for Energy-Focused Transportation Science

The Energy Efficient Mobility Systems (EEMS) technology landscape is complex and rapidly evolving, which provides both tremendous opportunities and formidable challenges. Significant alterations to the mobility landscape are underway due to the advent of vehicle and infrastructure connectivity, autonomous driving, and rapid passenger- and freight-vehicle electrification. Advanced computing will play an increasingly important role in enabling the EEMS program to understand and identify the most important levers to improve the energy productivity of future integrated mobility systems. It is also driving new approaches to mobility and the research to unlock an affordable, efficient, safe, and accessible transportation future. Driving much of this change is the collection, analysis, and strategic use of massive amounts of diverse, complex data from infrastructure and vehicles with on-board sensors and data storage and transmission capabilities. Diverse and representative data are key to implementing approaches to maximize mobility energy productivity. While high-fidelity modeling of integrated transportation networks has strengthened our understanding of dynamic movement and behavior patterns, existing tools must be expanded beyond their current focus. This work necessitates data infrastructure investments (e.g., secure-streaming data platforms driven by ubiquitous sensors and video analytics) as well as investments in critical capabilities for large-scale automated analysis and organization using modern machine learning, statistics, and artificial intelligence. Other chief needs include agile, large-scale storage that can be quickly searched and queried for relevant data to support validation and model development, data-sharing agreements, and formatting standards for key data types. The future of public transit must be explored in greater detail, research must inform design, and opportunities must be identified for improving the mobility productivity of public transit in both urban and rural America.

33 ADVANCED PROPULSION SYSTEMS↗

Mesoscale Science Data Analytics

This is software that will be used to do data analytics in experimental workflows for x-ray mesoscale science. This tool set will provide a mechanism for supporting experiments in many ways, from collecting calibration information and raw data, to managing and viewing data to extracting crystallographic and physical parameters. It will eventually include development of a fully automated workflow that will include statistical information and prediction capabilities to support the scientists in decision-making and replanning their experiments when necessary.

Sweeney, Christine↗

On the integration of molecular dynamics, data science, and experiments for studying solvent effects on catalysis

Computational workflows that combine molecular dynamics (MD) simulations and emerging data-centric (DC) methods can accelerate the screening and analysis of solvent systems experimentally and computationally. Here, MD simulations provide atomic positions and velocities of reactant, solvent, and catalyst materials that can be manipulated into data representations that in turn can be used by DC techniques to conduct predictive modeling, feature extraction, and experimental design. For liquid-phase catalytic applications, emerging DC techniques such as Convolutional and Graph Neural Networks (CNN/GNN), Topological Data Analysis (TDA), and Active Learning (AL) can leverage MD and experimental data to quickly predict solvent effects on reaction outcomes. For instance, in recent studies, 3D solvent environments obtained with MD have been exploited by CNNs to predict experimental reaction rates for homogeneous acid-catalyzed lignocellulosic processes. In this perspective, we discuss basic principles of DC methods and how these can be combined with MD to enable high-throughput screening of solvent selection for diverse catalysis applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Towards A Geo-Data Science Method for Assessing Rare Earth Element and Critical Mineral Occurrences in Coal and Other Sedimentary Systems

While preliminary analyses of data from open-source resources (Ekmann, 2012) show promising concentrations of REE in individual coal samples from a number of sites and basins in the U.S., other sparse data for REE in domestic coal-related strata suggest that many occurrences are low, “subeconomic” concentrations. At present, there is no method for systematically assessing potential sedimentary occurrences of REE. However, the geologic processes responsible for REE occurrences in coal-related strata are systematic; the unpredictability of REE resources in coal-related strata is due to poorly quantified spatial resource trends and the lack of an exploration method tailored to these resources. Thus, there exists a need for a systematic assessment approach that incorporates knowledge of geological variation in the mechanisms of REE enrichment within coal basins to help minimize geologic uncertainty and reduce commercial exploration risk.

01 COAL, LIGNITE, AND PEAT↗

A data science approach for analysis and reconstruction of spinodal-like composition fields in irradiated FeCrAl alloys

A statistical method for the analysis of continuously distributed data representative of composition fluctuations in irradiated FeCrAl alloys acquired using Energy Dispersive X-ray Spectroscopy (EDS) method is presented. Using probability distribution functions, direct and cross-covariances between the elemental compositions, the effects of alloy composition and irradiation dose were investigated on the spatial distribution and length scale of composition fluctuations at the nanoscale. We have observed that, for neutron-irradiated FeCrAl alloys, the distribution of Fe and Cr followed a left-skewed and right-skewed distribution, respectively for all (average) alloy compositions and irradiation doses. The analysis also revealed enhanced spatial gradients in the elemental compositions at higher irradiation dose. Direct and cross-covariance estimates of the experimental data were also utilized for reconstruction of composition data through fitting it to a parametric form of the covariance functions. Linear Model of Coregionalization was used to determine the parameters of the covariance functions. Subsequently, a spectral method was utilized for simulating a realization of the alloy compositions. Close correspondence was observed between the experimental and the reconstructed data which was analyzed using probability distribution functions and covariance functions. Composition space of the experimental and reconstructed data and dislocation velocities as a function of applied stress and line directions over the entire composition maps were also examined.

36 MATERIALS SCIENCE↗