Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Department of Energy’s Atmospheric System Research (ASR) Program’s Workshop on the Future of Atmospheric Large Eddy Simulation (LES) (Workshop Report)

Large-eddy simulation (LES) is used as a tool to understand physical processes such as turbulence, aerosols, clouds, precipitation, radiation, the interactions among all these, and their interactions with the underlying surface. Over the next 10 years, LES will drive fundamental progress in open scientific questions in these areas as LES is increasingly used to gain understanding of complex interacting physical processes involving atmospheric turbulence. This growth will be driven both by scientific demand and the expansion of computational resources needed to conduct LES, and the form that the growth takes will largely be determined by how computational resources are leveraged for scientific gain. In particular, we suggest that computational resources are likely to be leveraged in two separate but not necessarily distinct ways. On one hand, growth in computational resources will allow LES to be made more routine, that is, performed more frequently, while on the other hand, the computational expense (measured in total floating point operations) afforded to individual LES will expand dramatically, allowing simulations to increase in both domain size and resolution as well as physical detail. Current U.S. Department of Energy (DOE) projects such as LES ARM Symbiotic Simulation and Observation Activity (LASSO) are leading the way in conducting routine LES, building large, public databases that are accessible for data science, sensitivity studies, and training for machine learning. LES will also become more routine as it becomes more accessible for individual researchers to address their scientific questions of interest. Scientific questions addressed by LES over the next 10 years are likely to include cloud organization and aggregation; aerosol cloud interactions and atmospheric chemistry (including geo-engineering); urban-scale LES; atmospheric extreme events, ranging from small-scale severe weather to wildfires; and ocean-wave-atmosphere interactions. Further LES-related research will likely grow significantly in areas related to societal impact studies of air quality and extreme weather events, applications to renewable energy forecasting and resource assessment, and aid in decision-making processes.

54 ENVIRONMENTAL SCIENCES↗

Advanced Offshore Hazard Forecasting to Enable Resilient Offshore Operations

Paper prepared for the Offshore Technology Conference, 2024. Hazards in the offshore environment can imperil successful energy operations, whether those operations are conventional, renewable, or for decarbonization. The expanding accessibility of data science and the advanced applications of machine learning (ML) models creates an opportunity to assess potential hazards and the infrastructure they impact. We present a use case demonstrating the combined application of published ML tools to U.S. federal waters of the Gulf of Mexico, an actively explored region for offshore energy that is affected by variable metocean conditions and geologic processes contributing to potential hazards.

Mark-Moser, Mackenzie K.↗

Towards Trust-Augmented Visual Analytics for Data-Driven Energy Modeling

The promise of data-driven predictive modeling is being increasingly realized in various science and engineering disciplines, where experts are used to the more conventional, simulation-driven modeling practices. However, trust remains a bottleneck for greater adoption of machine learning-based models for domain experts, who might not be necessarily trained in data science. In this paper, we focus on the building energy domain, where physics-based simulations are being complemented or replaced by machine learning-based methods for forecasting energy supply and demand at various spatio-temporal scales. We study the trust problem in close collaboration with energy scientists and engineers and describe how visual analytics can be leveraged for alleviating this trust bottleneck for stakeholders with varying degrees of expertise and analytics goals in this domain.

Kandakatla, Akshith R.↗

Rapid data acquisition and machine learning-assisted composition design of functionally graded alloys via wire arc additive manufacturing

Abstract The lack of high-quality datasets in materials science hinders artificial intelligence (AI)-driven alloy design. To address this challenge, wire arc additive manufacturing (WAAM) was employed to fabricate graded alloys, generating extensive data for machine learning (ML)-assisted property prediction. ML models were developed using high-throughput experiments, computational models, and genetic algorithm to optimize feature selection, successfully predicting hardness and porosity. The ML model demonstrated its efficacy by designing a gradient alloy with enhanced properties. However, scaling up revealed uncertainties in tensile property and porosity due to differences in size and thermal conditions between the designed alloy build and the gradient print used to construct the ML model. This underscores the need for uncertainty quantification and process optimization in WAAM-driven alloy design. Our work advances AI-integrated additive manufacturing, offering a rapid approach to exploring process–structure–property relationships and accelerating materials development.

Wang, Xin↗

ASCENDS: Advanced data SCiENce toolkit for Non-Data Scientists

Recently, advances in machine learning and artificial intelligence have been playing more and more critical roles in a wide range of areas. For the last several years, industries have shown that how learning from data, identifying patterns and making decisions with minimal human intervention can be extremely useful to their business (e.g., image classification, recommending a product to a customer, finding friends in a social network, predicting customers actions, etc.). These success stories have been motivating scientists who study physics, chemistry, materials, medicine, and many other subjects, to explore a new pathway of utilizing machine learning techniques like regression and classification for their scientific activities. However, most existing machine learning tools, systems, and methodologies have been developed for programming experts but not for scientists (or any users) who have no or little knowledge of programming. ASCENDS is a toolkit that is developed to assist scientists (or any persons) who want to use their data for machine learning tasks, more specifically, correlation analysis, regression, and classification. ASCENDS does not require programming skills. Instead, it provides a set of simple but powerful CLI (Command Line Interface) and GUI (Graphic User Interface) tools for non-data scientists to be able to intuitively perform advanced data analysis and machine learning techniques. ASCENDS has been implemented by wrapping around opensource software including Keras, TensorFlow, and scikit-learn.

97 MATHEMATICS AND COMPUTING↗

Combustion machine learning: Principles, progress and prospects

Progress in combustion science and engineering has led to the generation of large amounts of data from large-scale simulations, high-resolution experiments, and sensors. This corpus of data offers enormous opportunities for extracting new knowledge and insights—if harnessed effectively. Machine learning (ML) techniques have demonstrated remarkable success in data analytics, thus offering a new paradigm for data-intense analyses and scientific investigations through combustion machine learning (CombML). While data-driven methods are utilized in various combustion areas, recent advances in algorithmic developments, the accessibility of open-source software libraries, the availability of computational resources, and the abundance of data have together rendered ML techniques ubiquitous in scientific analysis and engineering. This article examines ML techniques for applications in combustion science and engineering. Starting with a review of sources of data, data-driven techniques, and concepts, we examine supervised, unsupervised, and semi-supervised ML methods. Various combustion examples are considered to illustrate and to evaluate these methods. Next, we review past and recent applications of ML approaches to problems in combustion, spanning fundamental combustion investigations, propulsion and energy-conversion systems, and fire and explosion hazards. Challenges unique to CombML are discussed and further opportunities are identified, focusing on interpretability, uncertainty quantification, robustness, consistency, creation and curation of benchmark data, and the augmentation of ML methods with prior combustion-domain knowledge.

33 ADVANCED PROPULSION SYSTEMS↗

2020 ETI Annual Summer School: Data Science and Engineering

The Consortium for Enabling Technologies & Innovation (ETI) was established in 2019 to address emerging technologies within the context of nuclear nonproliferation. ETI creates a research and education environment to support cross-cutting technologies across three core disciplines: 1) computer and engineering science research specifically in a form of machine learning and high performance computing (HPC), 2) advanced manufacturing, and 3) nuclear detection technologies. For outreach and development, ETI hosted the first of three summer schools from August 24-28, 2020 with the theme of “Data Science and Engineering”. The school was hosted in an on-line format and had over 200 participants. The recorded content is available on-line as a resource for students. The summer school had four modules: 1) Fundamentals of data Applications, 2) Computational Machine Learning, 3) Bayesian Modeling and Inference, and 4) Data Science for Safeguards. Modules contained both lectures as well as student exercises. Poll Everywhere was utilized in some modules as an on-line method to engage large groups of students. Upcoming ETI Summer Schools include Novel Instrumentation in 2021 and Advanced Manufacturing in 2022.

Biegalski, Steven R.↗

Sub-pilot-scale Production of High-Value Products from U.S. Coals

Investigators from the University of Utah, University of Wyoming and Marshall University pursued a program to study the conversion of raw coal to high-value products of carbon fiber and silicon carbide. Team members also developed an initial framework for a data portal that can incorporate laboratory data on coal processing and product quality, and also work with tools for machine learning for data analysis, data visualization and economic assessment. Experimental R&D efforts focused on the conversion of raw coal to coal tar and other byproducts, and the resulting tar intermediates were upgraded to form anisotropic and isotropic pitch materials. These pitch materials were produced from coal using both thermal (pyrolysis) and chemical (mild solvolysis liquefaction) decomposition of raw coal. Four different coals were studied: Utah bituminous coal (Sufco), Wyoming PRB coal (Black Thunder), Illinois bituminous coal (Illinois #6), and West Virginia bituminous coal (Flying Eagle). Both metallurgical-grade coking coals and lower-grade steam coals were investigated, and controlled secondary gas-phase reactions were used during a two-stage pyrolysis process to induce cracking and condensation reactions among the pyrolytic tar species. This approach successfully improved the performance of the lower grade coals for yielding pitch materials, with properties more consistent with a commercial-grade pitch that had previously demonstrated success for quality carbon fiber production. The use of waste plastic materials was also studied, to help improve physical and chemical characteristics of the intermediate tars and final pitch product; in particular, for lowering the pitch softening point to an acceptable level for melt spinning carbon fiber. Mild solvolysis liquefaction was also used as a method for producing pitch for carbon fiber production. As expected, significantly higher pitch yields were obtained using this approach, and waste plastic materials were also successfully used to reduce pitch softening point to an acceptable level. The plastic materials were also utilized to create a solvent for the mild solvolysis process, and this plastic-derived solvent was shown to provide results consistent with more expensive commercial chemical solvents, and could thus avoid the need for costly recovery and recycle of a liquefaction solvent. Additional experimental R&D focused on the production of silicon carbide (β-SiC) from the residual char byproduct from pitch production, and also on the production of carbon fiber from the anisotropic pitch. SiC was successfully synthesized using a mixture of residual char and sandstone at a ratio of 1:1. Reaction temperature and residence time were optimized and yielded a product purity of 81%. For carbon fiber production, the most successful pitch samples were obtained from the mild solvolysis liquefaction approach, combined with the use of a plastic (HDPE)-derived solvent. Fiber properties improved over time as laboratory fiber production methodologies improved, and final yields of carbon fiber were obtained with a diameter of 12.14 ± 1.10 um, Modulus of 173.73 ± 15.25 GPa, and Tensile Strength of 1.04 ± 0.10 GPa. A proof-of-concept Modern Community Research Data Portal (MCRDP) was developed and deployed for coal and coal-derived pitch characterization, with the full support of (i) remote web-based access, (ii) distributed analysis, (iii) interactive visualization and exploration, (iv) shared and long-term data access, (v) advanced query capabilities and (vi) real-time collaboration. The Coal to Products Data Portal “coaltoproducts.org” provides researchers with space to store and share data within a project, tools for analyzing and understanding data for scientific investigation, and the ability to publish data to the broader community for reproducibility. The portal leverages the Material Commons 2.0 (MC) platform developed by the Center for PRedictive Integrated Structural Materials Science (PRISMS) of the University of Michigan, to achieve long-term longevity of data collections and, more importantly, collaborative science. A number of data visualization tools were also assessed and implemented for interrogating the experimental and modeling data. The machine learning portion of this project analyzed datasets from two different coal conversion processes performed on a diverse set of coal samples from both the coal pyrolysis experiments and the solvent liquefaction experiments. The work was initiated by exploring standard regression models on the pyrolysis data, aiming to understand the impact of sample characteristics and processing conditions on key product metrics. Over the course of the project, the focus expanded to include a variety of machine learning tools, delving into both supervised and unsupervised learning methods. Models tested on the pyrolysis data included linear, ridge, lasso, elastic-net, Gaussian process, random forest regression, and AutoSklearn, and the approach was continually refined to enhance predictive accuracy and model interpretability. Similar techniques were applied to the liquefaction data with an additional focus on feature engineering. Along with mesophase content, additional outputs of interest were the pitch yield, softening point, and QI content. Insights derived from these analyses are crucial in determining the factors influencing the quality and yield of coal-derived products. As the work progressed, the research evolved from foundational model comparisons to analyses of random forests, decision paths, and feature importance scores. A thorough market analysis was performed to examine the prospects of coal-based carbon fibers. The best opportunities for coal come from its lower and more stable price relative to petroleum, particularly for subbituminous coals, which is the primary advantage that a coal refinery may have over a petroleum refinery. Before a commercial CTP production facility can be modeled, however, several things need to be understood regarding the nature of the would-be coal refinery. These include the technology to be deployed, the size of facility, the volume(s) of co-product(s), and the waste and emissions profile of the plant. The volume of co-products and waste may be substantial and will require separate market analysis to ensure viability. In the near-term, the importance of coal tar pitch, in the form of carbon pitch, to the aluminum and steel industries is likely to overshadow the alternative use of this material as an input for carbon fiber. The importance of steel and aluminum in building materials, and the need for carbon materials in their manufacturing, will ensure that demand for these products remains for the long run. In addition, carbon fiber may also be the best substitute for steel and aluminum well into the future. While society will eventually be able to shift production of much of its electricity needs to renewables, it will not be able to shift away from fossil fuels for production of high-strength construction and vehicular materials. Demand for carbon fiber is expected to increase quickly, but the volume of carbon fiber and the amount of coal that would be needed to produce even a sizeable share of this market may still be relatively small compared to current coal production. Thus, other coal-based products like graphene, graphite, carbon foams, resins, and carbon-based building products will play important roles in sustaining coal production as coal-fired power generation continues to decline.

01 COAL, LIGNITE, AND PEAT↗

Plant science decadal vision 2020–2030: Reimagining the potential of plants for a healthy and sustainable future

Abstract Plants, and the biological systems around them, are key to the future health of the planet and its inhabitants. The Plant Science Decadal Vision 2020–2030 frames our ability to perform vital and far‐reaching research in plant systems sciences, essential to how we value participants and apply emerging technologies. We outline a comprehensive vision for addressing some of our most pressing global problems through discovery, practical applications, and education. The Decadal Vision was developed by the participants at the Plant Summit 2019, a community event organized by the Plant Science Research Network. The Decadal Vision describes a holistic vision for the next decade of plant science that blends recommendations for research, people, and technology. Going beyond discoveries and applications, we, the plant science community, must implement bold, innovative changes to research cultures and training paradigms in this era of automation, virtualization, and the looming shadow of climate change. Our vision and hopes for the next decade are encapsulated in the phrase reimagining the potential of plants for a healthy and sustainable future. The Decadal Vision recognizes the vital intersection of human and scientific elements and demands an integrated implementation of strategies for research (Goals 1–4), people (Goals 5 and 6), and technology (Goals 7 and 8). This report is intended to help inspire and guide the research community, scientific societies, federal funding agencies, private philanthropies, corporations, educators, entrepreneurs, and early career researchers over the next 10 years. The research encompass experimental and computational approaches to understanding and predicting ecosystem behavior; novel production systems for food, feed, and fiber with greater crop diversity, efficiency, productivity, and resilience that improve ecosystem health; approaches to realize the potential for advances in nutrition, discovery and engineering of plant‐based medicines, and "green infrastructure." Launching the Transparent Plant will use experimental and computational approaches to break down the phytobiome into a "parts store" that supports tinkering and supports query, prediction, and rapid‐response problem solving. Equity, diversity, and inclusion are indispensable cornerstones of realizing our vision. We make recommendations around funding and systems that support customized professional development. Plant systems are frequently taken for granted therefore we make recommendations to improve plant awareness and community science programs to increase understanding of scientific research. We prioritize emerging technologies, focusing on non‐invasive imaging, sensors, and plug‐and‐play portable lab technologies, coupled with enabling computational advances. Plant systems science will benefit from data management and future advances in automation, machine learning, natural language processing, and artificial intelligence‐assisted data integration, pattern identification, and decision making. Implementation of this vision will transform plant systems science and ripple outwards through society and across the globe. Beyond deepening our biological understanding, we envision entirely new applications. We further anticipate a wave of diversification of plant systems practitioners while stimulating community engagement, underpinning increasing entrepreneurship. This surge of engagement and knowledge will help satisfy and stoke people's natural curiosity about the future, and their desire to prepare for it, as they seek fuller information about food, health, climate and ecological systems.

59 BASIC BIOLOGICAL SCIENCES↗

When physics-informed data analytics outperforms black-box machine learning: A case study in thickness control for additive manufacturing

Aerosol jet printing (AJP) has emerged as a promising noncontact additive manufacturing method for high-resolution printing for a wide range of material systems. A key challenge limiting the broader adoption of AJP in the material science community is the lack of methods to precisely control thickness. Herein, we develop a model-based design of experiment (MBDoE) framework that integrates physics-informed models, nonlinear regression, and information criteria to postulate, select and calibrate the best model to describe and optimize the AJP manufacturing process. Starting with already available data from system commissioning (e.g., prior single variable sensitivity analysis), four candidate physics-informed models are postulated and trained. MBDoE identifies a single additional optimal experiment to validate these predictive models with quantified uncertainties, which are then used to determine the best experimental conditions to control printed film thickness. As a comparative benchmark, the analysis is repeated using the same dataset with nonparametric Gaussian process regression (GPR) model that does not incorporate physical information. Using MBDoE principles, we find that only five experiments are necessary to calibrate the nonlinear physics-informed parametric model, and with said limited data, this model outperforms the black-box machine learning GPR model. This key result underscores an emerging trend in the data science community: incorporating physical information into predictive models often drastically reduces the data requirements. Leveraging MBDoE further increased the data efficiency. By design, the proposed data science framework is general in nature and can be easily extended to other experimental and additive manufacturing systems beyond AJP.

Aerosol jet printing↗

Machine learning in nuclear materials research

Nuclear materials are often demanded to function for extended time in extreme environments, including high radiation fluxes with associated transmutations, high temperature and temperature gradients, mechanical stresses, and corrosive coolants. They also have a wide range of microstructural and chemical makeups, resulting in multifaceted and often out-of-equilibrium interactions. Machine learning (ML) is increasingly being used to tackle these complex time-dependent interactions and aid researchers in developing models and making predictions, sometimes with better accuracy than traditional modeling that focuses on one or two parameters at a time. Conventional practices of acquiring new experimental data in nuclear materials research are often slow and expensive, limiting the opportunity for data-centric ML, but new methods are changing that paradigm. Here we review high-throughput computational and experimental data approaches, especially robotic experimentation and active learning that is based on Gaussian process and Bayesian optimization. We show ML examples in structural materials (e.g., reactor pressure vessel (RPV) alloys and radiation detecting scintillating materials) and highlight new techniques of high-throughput sample preparation and characterizations, and automated radiation/environmental exposures and real-time online diagnostics. Herein, this review suggests that ML models of material constitutive relations in plasticity, damage, and even electronic and optical responses to radiation are likely to become powerful tools as they develop. Finally, we speculate on how the recent trends of using natural language processing (NLP) to aid the collection and analysis of literature data, interpretable artificial intelligence (AI), and the use of streamlined scripting, database, workflow management, and cloud computing platforms that will soon make the utilization of ML techniques as commonplace as the spreadsheet curve-fitting practices of today.

36 MATERIALS SCIENCE↗

OLCF’s Advanced Computing Ecosystem (ACE): FY25 Update for Ongoing Efforts

The advent of widespread use of artificial intelligence (AI) and machine learning (ML) models in science, coupled with fast data production rates of scientific instruments strain the traditional batch-oriented high-performance computing (HPC) environment. As scientific exploration continues to require more data and faster processing and analysis, new emerging technologies and capabilities to enable cross-facility and time-sensitive workflows are required for seamless integration of HPC and experimental facilities. The Advanced Computing Ecosystem (ACE) is a strategic initiative within the Oak Ridge Leadership Computing Facility (OLCF) established in 2024 to support the development of cutting-edge technologies to advance computational research and infrastructure at OLCF and across the Department of Energy (DOE). Several DOE initiatives are spearheading the evolution of the scientific landscape by blurring facility boundaries and connecting the user facilities to advance scientific capabilities and ensure energy dominance. The DOE Integrated Research Infrastructure (IRI) program is one example that is laying a foundation to support complex cross-facility workflows. The IRI program aims to integrate diverse computational resources, data infrastructures, and scientific instruments to facilitate collaboration and accelerate scientific discovery. The Interconnected Science Ecosystem (INTERSECT) initiative at Oak Ridge National Laboratory (ORNL) is another example that aims to revolutionize scientific research through AI-driven, interconnected autonomous laboratories and research facilities. Finally, the American Science Cloud (AmSC), recently announced in the “One Big Beautiful Bill”, aims to leverage prior infrastructure efforts of the IRI and automation and AI efforts of INTERSECT (and others) to build a federated, AI-augmented AmSC platform to unify the DOE’s computing, experimental, and data resources to catalyze scientific innovation.

97 MATHEMATICS AND COMPUTING↗

Efficient high-dimensional variational data assimilation with machine-learned reduced-order models

Abstract. Data assimilation (DA) in geophysical sciences remains the cornerstone of robust forecasts from numerical models. Indeed, DA plays a crucial role in the quality of numerical weather prediction and is a crucial building block that has allowed dramatic improvements in weather forecasting over the past few decades. DA is commonly framed in a variational setting, where one solves an optimization problem within a Bayesian formulation using raw model forecasts as a prior and observations as likelihood. This leads to a DA objective function that needs to be minimized, where the decision variables are the initial conditions specified to the model. In traditional DA, the forward model is numerically and computationally expensive. Here we replace the forward model with a low-dimensional, data-driven, and differentiable emulator. Consequently, gradients of our DA objective function with respect to the decision variables are obtained rapidly via automatic differentiation. We demonstrate our approach by performing an emulator-assisted DA forecast of geopotential height. Our results indicate that emulator-assisted DA is faster than traditional equation-based DA forecasts by 4 orders of magnitude, allowing computations to be performed on a workstation rather than a dedicated high-performance computer. In addition, we describe accuracy benefits of emulator-assisted DA when compared to simply using the emulator for forecasting (i.e., without DA). Our overall formulation is denoted AIEADA (Artificial Intelligence Emulator-Assisted Data Assimilation).

58 GEOSCIENCES↗

On the Use of Satellite Nightlights for Power Outages Prediction

Hurricanes are a dominant disaster in the Caribbean, always causing serious power outages throughout the islands. Hurricane Maria was a prime example, causing unimaginable destruction of the power infrastructure of Puerto Rico (PR). Consequently, one month after the hurricane landfall, approximately 80% of the population was still without power. After an event of such massive destruction, the electric power restoration process progresses very slowly. This timeline can be improved using power outage (PO) forecast models that help identify the vulnerable places before the hurricane landfall. Generally, these models are trained with historical power outages records, associated data on weather conditions, and additional information about the natural and built environments. However, PO records are often difficult to acquire, and, in many instances, the power utility companies may not record them. This study utilizes a satellite-based Visible Infrared Imaging Radiometer Suite (VIIRS) night light data product as a surrogate for the power delivery to predict hurricane-induced PO in areas having limited to nonexistent historical data records. The processed satellite data is then used along with geographic variables, and simulated weather data to formulate machine learning-based algorithms to predict PO for future hurricane events. These models are applied and validated in the context of the PR catastrophic storm, Hurricane Maria.

54 ENVIRONMENTAL SCIENCES↗