Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

HAM: Hotspot-Aware Manager for Improving Communications with 3D-Stacked Memory

merging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. %and an associated hotspot-aware page policy. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51\% and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81\% in average (up to 34.28\%) and power savings of 35.07\% over a standard 3D-stacked memory.

Wang, Xi↗

Generalized Canonical Polyadic Tensor Decomposition

Tensor decomposition is a fundamental unsupervised machine learning method in data science, with applications including network analysis and sensor data processing. This work develops a generalized canonical polyadic (GCP) low-rank tensor decomposition that allows other loss functions besides squared error. For instance, we can use logistic loss or Kullback--Leibler divergence, enabling tensor decomposition for binary or count data. We present a variety of statistically motivated loss functions for various scenarios. We provide a generalized framework for computing gradients and handling missing data that enables the use of standard optimization methods for fitting the model. Furthermore, we demonstrate the flexibility of the GCP decomposition on several real-world examples including interactions in a social network, neural activity in a mouse, and monthly rainfall measurements in India.

97 MATHEMATICS AND COMPUTING↗

Facilitating Machine Learning Collaborations Between Labs, Universities, And Industry

It is clear from numerous recent community reports, papers, and proposals that machine learning is of tremendous interest for particle accelerator applications. The quickly evolving landscape continues to grow in both the breadth and depth of applications including physics modeling, anomaly detection, controls, diagnostics, and analysis. Consequently, laboratories, universities, and companies across the globe have established dedicated machine learning (ML) and data science efforts aiming to make use of these new state-of-the-art tools. The current funding environment in the U.S. is structured in a way that supports specific application spaces rather than larger collaboration on community software. Here, we discuss the existing collaboration bottlenecks and how a shift in the funding environment, and how we develop collaborative tools, can help fuel the next wave of ML advancements for particle accelerators.

Edelen, J.P.↗

Identifying COVID-19 cases and extracting patient reported symptoms from Reddit using natural language processing

We used social media data from “covid19positive” subreddit, from 03/2020 to 03/2022 to identify COVID-19 cases and extract their reported symptoms automatically using natural language processing (NLP). We trained a Bidirectional Encoder Representations from Transformers classification model with chunking to identify COVID-19 cases; also, we developed a novel QuadArm model, which incorporates Question-answering, dual-corpus expansion, Adaptive rotation clustering, and mapping, to extract symptoms. Our classification model achieved a 91.2% accuracy for the early period (03/2020-05/2020) and was applied to the Delta (07/2021–09/2021) and Omicron (12/2021–03/2022) periods for case identification. We identified 310, 8794, and 12,094 COVID-positive authors in the three periods, respectively. The top five common symptoms extracted in the early period were coughing (57%), fever (55%), loss of sense of smell (41%), headache (40%), and sore throat (40%). During the Delta period, these symptoms remained as the top five symptoms with percent authors reporting symptoms reduced to half or fewer than the early period. During the Omicron period, loss of sense of smell was reported less while sore throat was reported more. Our study demonstrated that NLP can be used to identify COVID-19 cases accurately and extracted symptoms efficiently.

60 APPLIED LIFE SCIENCES↗

Veridical data science

Building and expanding on principles of statistics, machine learning, and scientific inquiry, we propose the predictability, computability, and stability (PCS) framework for veridical data science. Our framework, composed of both a workflow and documentation, aims to provide responsible, reliable, reproducible, and transparent results across the data science life cycle. The PCS workflow uses predictability as a reality check and considers the importance of computation in data collection/storage and algorithm design. It augments predictability and computability with an overarching stability principle. Stability expands on statistical uncertainty considerations to assess how human judgment calls impact data results through data and model/algorithm perturbations. As part of the PCS workflow, we develop PCS inference procedures, namely PCS perturbation intervals and PCS hypothesis testing, to investigate the stability of data results relative to problem formulation, data cleaning, modeling decisions, and interpretations. We illustrate PCS inference through neuroscience and genomics projects of our own and others. Moreover, we demonstrate its favorable performance over existing methods in terms of receiver operating characteristic (ROC) curves in high-dimensional, sparse linear model simulations, including a wide range of misspecified models. Finally, we propose PCS documentation based on R Markdown or Jupyter Notebook, with publicly available, reproducible codes and narratives to back up human choices made throughout an analysis. The PCS workflow and documentation are demonstrated in a genomics case study available on Zenodo.

97 MATHEMATICS AND COMPUTING↗

Turbulence theories and statistical closure approaches

When discussing research in physics and in science more generally, it is common to ascribe equal importance to the three components of the scientific trinity: theoretical, experimental, and computational studies. This review will explore the future of modern turbulence theory by tracing its history, which began in earnest with Kolmogorov’s 1941 analysis of turbulence cascade and inertial range [A.N. Kolmogorov, Dokl. Akad. Nauk SSSR, 30, 299, (1941); 32, 19, (1941)]. The 80th Anniversary of Kolmogorov’s landmark study is a welcome opportunity to survey the achievements and evaluate the future of the theoretical approach of turbulence research. Over the years, turbulence theories have been critically important in laying the foundation of our understanding of the nature of turbulent flows. In particular, the Direct Interaction Approximation (DIA) [R.H. Kraichnan, J. Fluid Mech., 5, 497 (1959)] and its subsequent development, known as the statistical closure approach, can be identified as perhaps the most profound single advancement. The remarkable success of the statistical closure has furnished a platform to study such essential concepts as the energy transfer process and interacting scales, and the roles of the straining and sweeping motions. More recently, the quasi-Lagrangian formulation of V. L’vov & I. Procaccia and Kraichnan’s solvable passive scalar model provided powerful ways to explore another fundamental aspect of turbulent flows, the phenomena of intermittency, and the associated anomalous scaling exponents. In the meantime, the theory of fluid equilibria has been developed to describe the large-scale structures that can emerge from turbulent cascades of two-dimensional and geophysical flows at a later time. And yet, despite all these successes, analytical treatments suffer from mathematical complexities. As a result, the utility of theoretical approaches has been limited to relatively idealized flows. On the other hand, in recent decades, computational abilities and experimental facilities have reached an unprecedented scale. Looking beyond the horizon, the imminent deployment of exascale supercomputers will generate complete datasets of the entire flow field of key benchmark flows, allowing researchers to extract additional measurements concerning fully developed, complex turbulent flow fields far beyond those available from the statistical closure theories. Some other developments that could potentially influence the future course of turbulence theories include the advancement of machine learning, artificial intelligence, and data science; likely disruptions arising from the advent of quantum computation; and the increasingly prominent role of turbulence research in providing more accurate climate scientific data. Finally, turbulence theorists can leverage these developments by asking the right questions and developing advanced, sophisticated frameworks that will be able to predict and correlate vast amounts of data from the other two components of the trinity.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Modern Senicide in the Face of a Pandemic: An Examination of Public Discourse and Sentiment About Older Adults and COVID-19 Using Machine Learning

Objectives This study examined public discourse and sentiment regarding older adults and COVID-19 on social media and assessed the extent of ageism in public discourse. Methods Twitter data (N = 82,893) related to both older adults and COVID-19 and dated from January 23 to May 20, 2020, were analyzed. We used a combination of data science methods (including supervised machine learning, topic modeling, and sentiment analysis), qualitative thematic analysis, and conventional statistics. Results The most common category in the coded tweets was “personal opinions” (66.2%), followed by “informative” (24.7%), “jokes/ridicule” (4.8%), and “personal experiences” (4.3%). The daily average of ageist content was 18%, with the highest of 52.8% on March 11, 2020. Specifically, more than 1 in 10 (11.5%) tweets implied that the life of older adults is less valuable or downplayed the pandemic because it mostly harms older adults. A small proportion (4.6%) explicitly supported the idea of just isolating older adults. Almost three-quarters (72.9%) within “jokes/ridicule” targeted older adults, half of which were “death jokes.” Also, 14 themes were extracted, such as perceptions of lockdown and risk. A bivariate Granger causality test suggested that informative tweets regarding at-risk populations increased the prevalence of tweets that downplayed the pandemic. Discussion Ageist content in the context of COVID-19 was prevalent on Twitter. Information about COVID-19 on Twitter influenced public perceptions of risk and acceptable ways of controlling the pandemic. Finaly, public education on the risk of severe illness is needed to correct misperceptions.

60 APPLIED LIFE SCIENCES↗

Tutorial: Lessons Learned for Behavior Analysts from Data Scientists

Big data is a computing term used to refer to large and complex data sets, typically consisting of terabytes or more of diverse data that is produced rapidly. The analysis of such complex data sets requires advanced analysis techniques with the capacity to identify patterns and abstract meanings from the vast data. The field of data science combines computer science with mathematics/statistics and leverages artificial intelligence, in particular machine learning, to analyze big data. This field holds great promise for behavior analysis, where both clinical and research studies produce large volumes of diverse data at a rapid pace (i.e., big data). This article presents basic lessons for the behavior analytic researchers and clinicians regarding integration of data science into the field of behavior analysis. We provide guidance on how to collect, protect, and process the data, while highlighting the importance of collaborating with data scientists to select a proper machine learning model that aligns with the project goals and develop models with input from human experts. Here, we hope this serves as a guide to support the behavior analysts interested in the field of data science to advance their practice or research, and helps them avoid some common pitfalls.

42 ENGINEERING↗

A representation-independent electronic charge density database for crystalline materials

In addition to being the core quantity in density functional theory, the charge density can be used in many tertiary analyses in materials sciences from bonding to assigning charge to specific atoms. The charge density is data-rich since it contains information about all the electrons in the system. With increasing utilization of machine-learning tools in materials sciences, a data-rich object like the charge density can be utilized in a wide range of applications. The database presented here provides a modern and user-friendly interface for a large and continuously updated collection of charge densities as part of the Materials Project. In addition to the charge density data, we provide the theory and code for changing the representation of the charge density which should enable more advanced machine-learning studies for the broader community.

36 MATERIALS SCIENCE↗

A representation-independent electronic charge density database for crystalline materials

Abstract In addition to being the core quantity in density-functional theory, the charge density can be used in many tertiary analyses in materials sciences from bonding to assigning charge to specific atoms. The charge density is data-rich since it contains information about all the electrons in the system. With the increasing prevalence of machine-learning tools in materials sciences, a data-rich object like the charge density can be utilized in a wide range of applications. The database presented here provides a modern and user-friendly interface for a large and continuously updated collection of charge densities as part of the Materials Project. In addition to the charge density data, we provide the theory and code for changing the representation of the charge density which should enable more advanced machine-learning studies for the broader community.

36 MATERIALS SCIENCE↗

Multimodal X-ray nano-spectromicroscopy analysis of chemically heterogeneous systems

Abstract Understanding the nanoscale chemical speciation of heterogeneous systems in their native environment is critical for several disciplines such as life and environmental sciences, biogeochemistry, and materials science. Synchrotron-based X-ray spectromicroscopy tools are widely used to understand the chemistry and morphology of complex material systems owing to their high penetration depth and sensitivity. The multidimensional (4D+) structure of spectromicroscopy data poses visualization and data-reduction challenges. This paper reports the strategies for the visualization and analysis of spectromicroscopy data. We created a new graphical user interface and data analysis platform named XMIDAS (X-ray multimodal image data analysis software) to visualize spectromicroscopy data from both image and spectrum representations. The interactive data analysis toolkit combined conventional analysis methods with well-established machine learning classification algorithms (e.g. nonnegative matrix factorization) for data reduction. The data visualization and analysis methodologies were then defined and optimized using a model particle aggregate with known chemical composition. Nanoprobe-based X-ray fluorescence (nano-XRF) and X-ray absorption near edge structure (nano-XANES) spectromicroscopy techniques were used to probe elemental and chemical state information of the aggregate sample. We illustrated the complete chemical speciation methodology of the model particle by using XMIDAS. Next, we demonstrated the application of this approach in detecting and characterizing nanoparticles associated with alveolar macrophages. Our multimodal approach combining nano-XRF, nano-XANES, and differential phase-contrast imaging efficiently visualizes the chemistry of localized nanostructure with the morphology. We believe that the optimized data-reduction strategies and tool development will facilitate the analysis of complex biological and environmental samples using X-ray spectromicroscopy techniques.

36 MATERIALS SCIENCE↗

Explaining and predicting human behavior and social dynamics in simulated virtual worlds: reproducibility, generalizability, and robustness of causal discovery methods

Ground Truth program was designed to evaluate social science modeling approaches using simulation test beds with ground truth intentionally and systematically embedded to understand and model complex Human Domain systems and their dynamics Lazer et al. (Science 369:1060–1062, 2020). Our multidisciplinary team of data scientists, statisticians, experts in Artificial Intelligence (AI) and visual analytics had a unique role on the program to investigate accuracy, reproducibility, generalizability, and robustness of the state-of-the-art (SOTA) causal structure learning approaches applied to fully observed and sampled simulated data across virtual worlds. In addition, we analyzed the feasibility of using machine learning models to predict future social behavior with and without causal knowledge explicitly embedded. In this paper, we first present our causal modeling approach to discover the causal structure of four virtual worlds produced by the simulation teams—Urban Life, Financial Governance, Disaster and Geopolitical Conflict. Our approach adapts the state-of-the-art causal discovery (including ensemble models), machine learning, data analytics, and visualization techniques to allow a human-machine team to reverse-engineer the true causal relations from sampled and fully observed data. We next present our reproducibility analysis of two research methods team’s performance using a range of causal discovery models applied to both sampled and fully observed data, and analyze their effectiveness and limitations. We further investigate the generalizability and robustness to sampling of the SOTA causal discovery approaches on additional simulated datasets with known ground truth. Our results reveal the limitations of existing causal modeling approaches when applied to large-scale, noisy, high-dimensional data with unobserved variables and unknown relationships between them. We show that the SOTA causal models explored in our experiments are not designed to take advantage from vasts amounts of data and have difficulty recovering ground truth when latent confounders are present; they do not generalize well across simulation scenarios and are not robust to sampling; they are vulnerable to data and modeling assumptions, and therefore, the results are hard to reproduce. Finally, when we outline lessons learned and provide recommendations to improve models for causal discovery and prediction of human social behavior from observational data, we highlight the importance of learning data to knowledge representations or transformations to improve causal discovery and describe the benefit of causal feature selection for predictive and prescriptive modeling.

97 MATHEMATICS AND COMPUTING↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Accelerating manufacturing for biomass conversion via integrated process and bench digitalization: a perspective

We present a perspective for accelerating biomass manufacturing via digitalization. We summarize the challenges for manufacturing and identify areas where digitalization can help. A profound potential in using lignocellulosic biomass and renewable feedstocks, in general, is to produce new molecules and products with unmatched properties that have no analog in traditional refineries. Discovering such performance-advantaged molecules and the paths and processes to make them rapidly and systematically can transform manufacturing practices. Furthermore, we discuss retrosynthetic approaches, text mining, natural language processing, and modern machine learning methods to enable digitalization. Laboratory and multiscale computation automation via active learning are crucial to complement existing literature and expedite discovery and valuable data collection without a human in the loop. Such data can help process simulation and optimization select the most promising processes and molecules according to economic, environmental, and societal metrics. We propose the close integration between bench and process scale models and data to exploit the low dimensionality of the data and transform the manufacturing for renewable feedstocks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-Driven Strategies for Accelerated Materials Design

The ongoing revolution of the natural sciences by the advent of machine learning and artificial intelligence sparked significant interest in the material science community in recent years. The intrinsically high dimensionality of the space of realizable materials makes traditional approaches ineffective for large-scale explorations. Modern data science and machine learning tools developed for increasingly complicated problems are an attractive alternative. An imminent climate catastrophe calls for a clean energy transformation by overhauling current technologies within only several years of possible action available. Tackling this crisis requires the development of new materials at an unprecedented pace and scale. For example, organic photovoltaics have the potential to replace existing silicon-based materials to a large extent and open up new fields of application. In recent years, organic light-emitting diodes have emerged as state-of-the-art technology for digital screens and portable devices and are enabling new applications with flexible displays. Reticular frameworks allow the atom-precise synthesis of nanomaterials and promise to revolutionize the field by the potential to realize multifunctional nanoparticles with applications from gas storage, gas separation, and electrochemical energy storage to nanomedicine. In the recent decade, significant advances in all these fields have been facilitated by the comprehensive application of simulation and machine learning for property prediction, property optimization, and chemical space exploration enabled by considerable advances in computing power and algorithmic efficiency. In this Account, we review the most recent contributions of our group in this thriving field of machine learning for material science. We start with a summary of the most important material classes our group has been involved in, focusing on small molecules as organic electronic materials and crystalline materials. Specifically, we highlight the data-driven approaches we employed to speed up discovery and derive material design strategies. Subsequently, our focus lies on the data-driven methodologies our group has developed and employed, elaborating on high-throughput virtual screening, inverse molecular design, Bayesian optimization, and supervised learning. We discuss the general ideas, their working principles, and their use cases with examples of successful implementations in data-driven material discovery and design efforts. Furthermore, we elaborate on potential pitfalls and remaining challenges of these methods. Finally, we provide a brief outlook for the field as we foresee increasing adaptation and implementation of large scale data-driven approaches in material discovery and design campaigns.

36 MATERIALS SCIENCE↗

Machine Learning to Select Experiments Driven by Fundamental Science and Applications for Targeted Nuclear Data Improvement

This work describes a blueprint for a process that accelerates progress in science by quantitatively answering the following question: What is the optimal combination of fundamental-science and application-driven experiments to maximally reduce pertinent data uncertainties? Answering this question entails solving a high-dimensional and complex optimization problem that is best solved with advanced statistic techniques often classified as machine learning. We apply this process within the framework of nuclear data with the aim to select an experiment combination that will reduce uncertainties in 239 Pu nuclear data for neutron energies between 1 and 600 keV. In this field, fundamental-physics driven data, called differential, look at one nuclear physics observable at a time. They are contrasted to application-driven, integral, data where one or few resulting values inform a broad set of nuclear data across several nuclides and energies. The candidates for integral experiments are criticality measurements that were refined by a genetic algorithm to be maximally sensitive to 239 Pu fission cross sections in the desired energy range. Twenty-three candidate differential experiments were investigated and span multiple nuclear physics observables (e.g., total, capture cross sections) for isotopes appearing in the integral experiments. The optimal combination among these candidate experiments was investigated via generalized least squares fitting, augmented with Gaussian processes to ameliorate statistical irregularities in data, and the D-optimality criterion. The latter evaluates for each pair of candidates the joint reduction in uncertainties of all 12200 nuclear data appearing in the integral experiments compared to the knowledge we have from 168 past experiments, theory, and nuclear data. We chose as differential measurements those that investigate 63 Cu and 239 Pu total cross sections, based on D-optimality rank and feasibility constraints. Two integral (criticality) experiments were selected: An experiment with Al 2 ⁢O 3 and graphite interleaved with Pu and a thick Cu reflector explores 1–30 keV, while we target the 30–600 keV range with an experiment that swaps boron in place of graphite with a different geometry.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Breeding Realistic D‐Brane Models

Abstract Intersecting branes provide a useful mechanism to construct particle physics models from string theory with a wide variety of desirable characteristics. The landscape of such models can be enormous, and navigating towards regions which are most phenomenologically interesting is potentially challenging. Machine learning techniques can be used to efficiently construct large numbers of consistent and phenomenologically desirable models. In this work we phrase the problem of finding consistent intersecting D‐brane models in terms of genetic algorithms, which mimic natural selection to evolve a population collectively towards optimal solutions. For a four‐dimensional supersymmetric type IIA orientifold with intersecting D6‐branes, we demonstrate that unique, fully consistent models can be easily constructed, and, by a judicious choice of search environment and hyper‐parameters, of the found models contain the desired Standard Model gauge group factor. Having a sizable sample allows us to draw some preliminary landscape statistics of intersecting brane models both with and without the restriction of having the Standard Model gauge factor.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Reinforced double-threaded slide-ring networks for accelerated hydrogel discovery and 3D printing

Traditionally, slide-ring gels are stretchable but soft as a result of an elasticity-stretchability trade-off. Herein, we introduce a new approach to breaking this trade-off and creating reinforced slide-ring networks with mobile crosslinkers. Our approach involves the construction of a polyethylene glycol double-threaded γ-cyclodextrin-based pro-slide-ring crosslinker that serves as a modular component for 3D printing and copolymerization. The resulting crystalline-domain-reinforced slide-ring hydrogels, or CrysDoS-gels, exhibit both high elasticity and high stretchability. The modular synthesis allows for high-throughput synthesis of CrysDoS-gels, generating a large amount of data for structure-property analysis. Here, by employing data science techniques, such as machine learning and linear regression, not only were we able to identify which chemical components influence the mechanical properties of CrysDoS-gels, but this analysis also aided in the discovery of better-performing CrysDoS-gels. Finally, we demonstrate the potential application of the newly discovered CrysDoS-gels as sensing devices by 3D printing them as stress sensors with high sensitivity and a broad detection range.

3D-printing↗