Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data- limited”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Modeling Electric Vehicle Charging Station Siting Suitability with a Focus on Equity

As adoption of electric vehicles increases, the infrastructure to charge them must keep pace. Determining where to add new charging infrastructure is a complex process subject to many factors, including electrical service availability, vehicle dwell time, the type(s) of drivers and vehicles the stations will serve, traffic levels and timing, and land ownership. In addition, advancing social equity is a current priority of federal efforts to invest in electric vehicle charging infrastructure. Conducting Multi-criteria Decision Analysis (MCDA) within Argonne’s Energy Zones Mapping Tool (EZMT) is a useful method for analyzing many of the factors that influence how suitable a location is for potentially adding new charging infrastructure, and we show how equity metrics can be included in the analysis. However, data limitations impose challenges to using MCDA to evaluate and prioritize locations. We use three examples to demonstrate how to use publicly available data and MDCA to analyze different siting objectives. Each example starts with defining a specific objective and ends with how to use the results to identify specific potential locations that could be investigated further. This analysis demonstrates how interested stakeholders can use the EZMT to run the example MCDA models defined in this study, modify them to suit their needs, or create new MCDA models.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Making a Water Data System Responsive to Information Needs of Decision Makers

Evidence-based environmental management requires data that are sufficient, accessible, useful and used. A mismatch between data, data systems, and data needs for decision making can result in inefficient and inequitable capital investments, resource allocations, environmental protection, hazard mitigation, and quality of life. In this paper, we examine the relationship between data and decision making in environmental management, with a focus on water management. We focus on the concept of decision-driven data systems —data systems that incorporate an assessment of decision-makers' data needs into their design. The aim of the research was to examine the process of translating data into effective decision making by engaging stakeholders in the development of a water data system. Using California's legislative mandate for state agencies to integrate existing water and other environmental data as a case study, we developed and applied a participatory approach to inform data-system design and identify unmet data needs. Using workshops and focused stakeholder meetings, we developed 20 diverse use cases to assess data sources, availability, characteristics, gaps, and other attributes of data used for representative decisions. Federal and state agencies made up about 90% of the data sources, and could readily adapt to a federated data system, our recommended model for the state. The remaining 10% of more-specialized data, central to important decisions across multiple use cases, would require additional investment or incentives to achieve data consistency, interoperability, and compatibility with a federated system. Based on this assessment, we propose a typology of different types of data limitations and gaps described by stakeholders. We also propose technical, governance, and stakeholder engagement evaluation criteria to guide planning and building environmental data systems. Data-system governance involving both producers and users of data was seen as essential to achieving workable standards, stable funding, convenient data availability, resilience to institutional change, and long-term buy-in by stakeholders. Our work provides a replicable lesson for using decision-maker and stakeholder engagement to shape the design of an environmental data system, and inform a technical design that addresses both user and producer needs.

Cantor, Alida↗

Efficient Reinforcement Learning for Real-Time Hardware-Based Energy System Experiments: Preprint

In the context of urgent climate challenges and the pressing need for rapid technology development, Reinforcement Learning (RL) stands as a compelling data-driven method for controlling real-world physical systems. However, RL implementation often entails time-consuming and computationally intensive data collection and training processes, rendering them inefficient for real-time applications that lack non-real-time models. To address these limitations, real-time emulation techniques have emerged as valuable tools for the lab-scale rapid prototyping of intricate energy systems. While emulated systems offer a bridge between simulation and reality, they too face constraints, hindering comprehensive characterization, testing, and development. In this research, we construct a surrogate model using limited data from simulated systems, enabling an efficient and effective training process for a Double Deep Q-Network (DDQN) agent for future deployment. Our approach is illustrated through a hydropower application, demonstrating the practical impact of our approach on climate-related technology development.

deep Q-learning↗

Scalable and Actionable Performance Measures for Traffic Signal Systems using Probe Vehicle Trajectory Data

Scalable and actionable performance measures for traffic signal systems provide opportunities for practitioners to measure and improve the transportation network. Historically, traffic signal improvements have relied on scheduled signal retiming based on limited data collection, or on the public to call and alert engineers of an issue. This inefficient method of improving signal timing led to the creation of automated traffic signal performance measures (ATSPMs). These metrics rely on expensive infrastructure, including detection and communications, which has produced barriers for numerous agencies to fully adopt. Recently, third-party data providers have begun to release vehicle trajectory data, which allows for enhanced signal metrics with no investment in physical equipment. The purpose of this study is to demonstrate the use of these data and summarize the scalability of the created metrics. This work builds on previous efforts to quantify signal performance on nine intersections in Michigan, U.S. Ten signalized corridors in Columbus, Ohio, were chosen to scale a performance assessment using crowdsourced trajectory data. A total of 136 intersections were assessed in 2-h intervals using data from all weekdays in 2017. High-level corridor summary metrics including average percent of vehicles stopping (18%–32%), average delay (9.4–20.5 s), and level of travel time reliability (1.23–2.73) were calculated for each corridor direction. Intersection-level metrics were also introduced, which can be used by practitioners to identify problems, improve signal timings, and prioritize future infrastructure investments.

99 GENERAL AND MISCELLANEOUS↗

Predicting the propensity for thermally activated β events in metallic glasses via interpretable machine learning

Abstract The elementary excitations in metallic glasses (MGs), i.e., β processes that involve hopping between nearby sub-basins, underlie many unusual properties of the amorphous alloys. A high-efficacy prediction of the propensity for those activated processes from solely the atomic positions, however, has remained a daunting challenge. Recently, employing well-designed site environment descriptors and machine learning (ML), notable progress has been made in predicting the propensity for stress-activated β processes (i.e., shear transformations) from the static structure. However, the complex tensorial stress field and direction-dependent activation could induce non-trivial noises in the data, limiting the accuracy of the structure-property mapping learned. Here, we focus on the thermally activated elementary excitations and generate high-quality data in several Cu-Zr MGs, allowing quantitative mapping of the potential energy landscape. After fingerprinting the atomic environment with short- and medium-range interstice distribution, ML can identify the atoms with strong resistance or high compliance to thermal activation, at a high accuracy over ML models for stress-driven activation events. Interestingly, a quantitative “between-task” transferring test reveals that our learnt model can also generalize to predict the propensity of shear transformation. Our dataset is potentially useful for benchmarking future ML models on structure-property relationships in MGs.

36 MATERIALS SCIENCE↗

Predicting Small Molecule Transfer Free Energies by Combining Molecular Dynamics Simulations and Deep Learning

Accurately predicting small molecule partitioning and hydrophobicity is critical in the drug discovery process. There are many heterogeneous chemical environments within a cell and entire human body. For example, drugs must be able to cross the hydrophobic cellular membrane to reach their intracellular targets, and hydrophobicity is an important driving force for drug–protein binding. Atomistic molecular dynamics (MD) simulations are routinely used to calculate free energies of small molecules binding to proteins, crossing lipid membranes, and solvation but are computationally expensive. Machine learning (ML) and empirical methods are also used throughout drug discovery but rely on experimental data, limiting the domain of applicability. We present atomistic MD simulations calculating 15,000 small molecule free energies of transfer from water to cyclohexane. This large data set is used to train ML models that predict the free energies of transfer. We show that a spatial graph neural network model achieves the highest accuracy, followed closely by a 3D-convolutional neural network, and shallow learning based on the chemical fingerprint is significantly less accurate. A mean absolute error of ~4 kJ/mol compared to the MD calculations was achieved for our best ML model. We also show that including data from the MD simulation improves the predictions, tests the transferability of each model to a diverse set of molecules, and show multitask learning improves the predictions. This work provides insight into the hydrophobicity of small molecules and ML cheminformatics modeling, and our data set will be useful for designing and testing future ML cheminformatics methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Reevaluation of Radiation-Protection Standards for Workers and the Public Based on Current Scientific Evidence

President Trump’s recent executive orders to “Usher in a Nuclear Renaissance,” coupled with the global pledge to triple nuclear energy capacity by 2050, underscore nuclear energy’s importance to national security and economic prosperity. This renewed interest has prompted efforts to spur nuclear-energy deployment, including assessing factors impeding it. One issue that has previously been identified as adding to the cost of nuclear energy is excessively conservative requirements, including those related to radiation protection. This technical review, therefore, examines current radiation-protection standards that were established decades ago when more limited data were available and nuclear-energy expansion was not a national priority. The review focuses on scientific evidence regarding the health effects of ionizing radiation at annual doses of 10,000 mrem or less. The review evaluates epidemiological studies, radiobiological research, and positions of relevant professional organizations to assess whether such doses result in discernable or observable increases in negative health outcomes. The review also surveys the literature related to economic and practical implications of current radiation-protection standards and practices. Based on this assessment, we propose maintaining an annual occupational whole-body dose limit of 5,000 mrem/yr and eliminating all “as low as reasonably achievable” requirements and limits below this threshold. This change could potentially reduce radiation-protection costs by millions of dollars annually for each reactor, as well as decrease the overall costs and correct misconceptions about the risks associated with all nuclear technologies. The evidence further supports future consideration of a 10,000 mrem/yr limit that would maintain appropriate safety margins while further reducing protection costs. Similarly, given the data and that the average annual radiation dose per person in the U.S. is 620 mrem, we believe the public dose limits of 100 mrem/yr are unnecessarily restrictive; increasing to 500 mrem/yr would maintain substantial safety margins—a factor of 10 below the occupational limit—while reducing regulatory burdens and associated bureaucracy. While we acknowledge ongoing scientific debate and encourage continued research on the health effects of ionizing radiation, our review indicates current frameworks are overly conservative. These overly stringent limits not only impose unnecessary economic burdens without corresponding health benefits but also divert safety focus and resources from more important considerations. Although this study was motivated by nuclear-power considerations, reforms to radiation-protection requirements have significant positive implications for other areas, such as nuclear medical applications, environmental remediation, nuclear-waste management and disposal, and industrial applications of nuclear technologies.

61 RADIATION PROTECTION AND DOSIMETRY↗

Correlated Trajectory Uncertainty for Adaptive Sequential Decision Making

One of the great challenges with decision making tasks on real world systems is the fact that data is sparse and acquiring additional data is expensive. In these cases, it is often crucial to make a model of the environment to assist in making decisions. At the same time, limited data means that learned models are erroneous, making it just as important to equip the model with good predictive uncertainties. In the context of learning sequential decision making policies, these uncertainties can prove useful for informing which data to collect for the greatest improvement in policy performance \citep{mehta2021experimental, mehta2022exploration} or informing the policy about unsure regions of state and action space to avoid during test time \citep{yu2020mopo}. Additionally, assuming that realistic samples of the environment can be drawn, an adaptable policy can be trained that attempts to make optimal decisions for any given possible instance of the environment \citep{ghosh2022offline, chen2021offline}. In this work, we examine the so-called ``probabilistic neural network'' (PNN) model that is ubiquitous in model-based reinforcement learning (MBRL) works. We argue that while PNN models may have good marginal uncertainties, they form a distribution of non-smooth transition functions. Not only are these samples unrealistic and may hamper adaptability, but we also assert that this leads to poor uncertainty estimates when predicting multiple step trajectory estimates. To address this issue, we propose a simple sampling method that can be implemented on top of pre-existing models.We evaluate our sampling technique on a number of environments, including a realistic nuclear fusion task, and find that, not only do smooth transition function samples produce more calibrated uncertainties, but they also lead to better downstream performance for an adaptive policy.

Offline Reinforcement Learning↗

AGR-5/6/7 Experiment Preliminary Findings

AGR-5/6/7 in-pile data have been published: Fuel irradiation conditions, Fission gas release-rate-to-birth-rate (R/B) ratios, PIE and safety testing are still in progress and only limited data have been published, and preliminary results.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Micrometer: Micromechanics transformer for predicting full field mechanical responses of heterogeneous materials

Predicting mechanical responses of heterogeneous materials across scales remains a significant challenge. Traditional computational methods often struggle with complex and multiscale nature of these materials, limiting their effectiveness in real-world applications. Here, in this paper, we introduce Micrometer, a vision transformer based deep learning model designed to predict full field mechanical responses of heterogeneous materials, bridging the gap between computer vision and solid mechanics problems. We show that Micrometer, trained on a large-scale high-resolution dataset of 2D fiber-reinforced composites, can achieve state-of-the-art performance in predicting microscale strain fields across a wide range of material properties and loading conditions. Our model demonstrates accuracy and computational efficiency in applications such as computational homogenization and multiscale modeling, reducing computational time by up to two orders of magnitude compared to conventional numerical solvers while maintaining less than 1 % errors in predicting macroscale stress fields. Furthermore, we showcase Micrometer’s adaptability through transfer learning experiments on new materials with limited data, highlighting its potential to tackle diverse scenarios in computational solid mechanics. These results represent a significant step towards AI-driven innovation in materials science, addressing the limitations of traditional numerical methods and paving the way for more efficient simulations of heterogeneous materials across various industrial applications.

Composite materials↗

The expressivity of classical and quantum neural networks on entanglement entropy

Abstract Analytically continuing the von Neumann entropy from Rényi entropies is a challenging task in quantum field theory. While then-th Rényi entropy can be computed using the replica method in the path integral representation of quantum field theory, the analytic continuation can only be achieved for some simple systems on a case-by-case basis. In this work, we propose a general framework to tackle this problem using classical and quantum neural networks with supervised learning. We begin by studying several examples with known von Neumann entropy, where the input data is generated by representing$${\text {Tr}}\rho _A^n$$ Tr ρ A n with a generating function. We adopt KerasTuner to determine the optimal network architecture and hyperparameters with limited data. In addition, we frame a similar problem in terms of quantum machine learning models, where the expressivity of the quantum models for the entanglement entropy as a partial Fourier series is established. Our proposed methods can accurately predict the von Neumann and Rényi entropies numerically, highlighting the potential of deep learning techniques for solving problems in quantum information theory.

Physics↗

A Survey of Constrained Gaussian Process: Approaches and Implementation Challenges

Gaussian process regression is a popular Bayesian framework for surrogate modeling of expensive data sources. As part of a larger effort in scientific machine learning, many recent works have incorporated physical constraints or other a priori information within Gaussian process regression to supplement limited data and regularize the behavior of the model. We provide an overview and survey of several classes of Gaussian process constraints, including positivity or bound constraints, monotonicity and convexity constraints, differential equation constraints provided by linear PDEs, and boundary condition constraints. We compare the strategies behind each approach as well as the differences in implementation, concluding with a discussion of the computational challenges introduced by constraints.

97 MATHEMATICS AND COMPUTING↗

Commercial, industrial, and institutional discount rate estimation for efficiency standards analysis: Sector-level data 1998–2021

Underlying each of the Department of Energy’s (DOE’s) federal appliance and equipment energy conservation standards are a set of complex analyses of the projected costs and benefits of regulation. Any new or amended standard must be designed to achieve significant additional energy conservation, provided that it is technologically feasible and economically justified (42 U.S.C. 6295(o)(2)(A)). DOE determines economic justification based on whether the benefits exceed the burdens, considering a variety of factors, including the economic impact of the standard on consumers of the product and the savings in lifetime operating cost compared to any increase in price or maintenance expenses (42 U.S.C. 6295(o)(2)(B)). As part of this determination, DOE conducts a Life-Cycle Cost (LCC) analysis, which models the combined impact of appliance first cost and operating cost changes on a representative commercial building sample in order to identify the fraction of customers achieving LCC savings or incurring net cost at the considered efficiency levels. Thus, the commercial discount rate value(s) used to calculate the present value of energy cost savings within the LCC model implicitly plays a role in estimating the economic impact of potential standard levels. This report provides an in-depth discussion of the commercial discount rate estimation process. It is an update to previous reports on estimating commercial discount rates from firm-level financial data (Fujita, 2016). Major topics covered in this report include: Discount rate estimation methods and rationale; -Data sources used and data limitations; -Discount rate distributions for use in standards analysis; -Discount rate estimation methods and distributions specific to the small business subgroup analysis. Going forward, this report will be updated as data allow and analyses necessitate.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Domain-aware Control-oriented Neural Models for Autonomous Underwater Vehicles

Conventional physics-based modeling is a time-consuming bottleneck in control design for complex nonlinear systems like autonomous underwater vehicles (AUVs). In contrast, purely data-driven models, require a large number of observations and lack operational guarantees for safety-critical systems. Data-driven models leveraging available partially characterized dynamics have potential to provide reliable systems models in a typical data-limited scenario for high value complex systems, thereby avoiding months of expensive expert modeling time. In this work we explore this middle-ground between expert-modeled and pure data-driven modeling. We present control-oriented parametric models with varying levels of domain-awareness that exploit known system structure and prior physics knowledge to create constrained deep neural dynamical system models. We employ universal differential equations to construct data-driven blackbox and graybox representations of the AUV dynamics. In addition, we explore a hybrid formulation that explicitly models the residual error related to imperfect graybox models. We compare the prediction performance of the learned models for different distributions of initial conditions and control inputs to assess their suitability for control.

Shaw Cortez, Wenceslao E.↗

Mixture Model for Refrigerant Pairs R-32/1234yf, R-32/1234ze(E), R-1234ze(E)/227ea, R-1234yf/152a, and R-125/1234yf

In this work, thermodynamic models based on the corresponding states framework with departure terms are developed for the refrigerant pairs R-32/1234yf, R-32/1234ze(E), R-1234ze(E)/227ea, R-1234yf/152a, and R-125/1234yf. These models are based on new measurements of density, speed of sound, and phase equilibria, combined with the data available in the literature. The model for R-32/1234yf is most comprehensive in its data coverage, with speed of sound deviations within 1%, density deviations within 0.1%, and bubble- and dew-point pressure deviations within 1%. In conclusion, the other mixtures have generally more limited data availability but a similar goodness of fit.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

East River Watershed Stable Water Isotope Data in Precipitation, Snowpack and Snowmelt 2016-2020

Stable water isotopes (d18O, d2H and d-excess) are important tracers in hydrologic research to understand water partitioning between vegetation, groundwater, and runoff but are rarely applied to large watersheds with persistent snowpack and complex topopgraphy. Data were collected for the Lawrence Berkeley National Laboratory Watershed Function Science Focus Area supported by the U.S Department of Energy in the East River, CO Hydrologic Unit Code (140200010) with limited data also collected in adjacent watersheds Ohio Creek and Taylor River. Data are provided in csv and includes isotopic information for precipitation (years 2014-2016), snowmelt (years 2016-2017) and snowpits (years 2016-2020). Snowpit data contain depth resolved information at 10 cm intervals for density, snow water equivalent (SWE) and stable water isotopes. Bulk isotopic data for 86 snowpits contain depth, SWE, density and SWE-weighted isotope values. Sampling locations and elevations are provided within the data files, while kmz files are provided to view sampling locations using Google Earth software.

54 ENVIRONMENTAL SCIENCES↗

A VOI Web Application for Distinct Geothermal Domains: Statistical Evaluation of Different Data Types within the Great Basin

The Great Basin region contains different domains that have different structural and hydrothermal flow patterns. Depending on the characteristics of these patterns, certain data types may be more successful at detecting hidden geothermal resources. In this paper, we quantitatively evaluate if certain data types are more successful in certain domains. Given different aquifer, strain and structural conditions, we explore which data types statistically reveal positively labeled geothermal sites. We utilize value of information (VOI) metrics to help quantify the reliability of data types to discriminate against "positive" and "negative" labeled geothermal sites. We also evaluate how kernel density estimation can help generalize the statistics that inform VOI, which is necessary given the limited data in geothermal exploration. Except for the Carbonate Aquifer, the highest ranking of the Vimperfect is the Local Structural Setting. Next, the slip and dilation tendency is first for Carbonate Aquifer and second for Central Nevada Seismic Belt and Western Great Basin. For the Carbonate Aquifer, heat flow is has the lowest Vimperfect value compared to the other three domains, which is consistent with the understanding of how heat flow measurements are masked by regional groundwater flow.

Bayesian analysis↗

Journey over Destination: Dynamic Sensor Placement Enhances Generalization

Reconstructing complex, high-dimensional global fields from limited data points is a challenge across various scientific and industrial domains. This is particularly important for recovering spatio-temporal fields using sensor data from, for example, laboratory-based scientific experiments, weather forecasting, or drone surveys. Given the prohibitive costs of specialized sensors and the inaccessibility of
certain regions of the domain, achieving full field coverage is typically not feasible. Therefore, the development of machine learning algorithms trained to reconstruct fields given a limited dataset is of critical importance. In this study, we introduce a general
approach that employs moving sensors to enhance data exploitation during the training of an attention based neural network, thereby improving field reconstruction. The training of sensor locations is accomplished using an end-to-end workflow, ensuring
differentiability in the interpolation of field values associated to the sensors, and is simple to implement using differentiable programming. Additionally, we have incorporated a correction mechanism to prevent sensors from entering invalid regions within the domain. We evaluated our method using two distinct datasets; the results show that our approach enhances learning, as evidenced by improved test scores.

54 ENVIRONMENTAL SCIENCES↗