Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data needs”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Report on the 4th Marine Energy Instrumentation and Data Workshop

The 4th Marine Energy Instrumentation and Data workshop was held on March 16 - 17 2022. This gathering brought together marine energy (ME) developers, researchers, and stakeholders to discuss the current state of ME technologies and the industry's instrument and data needs. The overall objective of the workshop was to identify gaps facing the ME industry for needed data collection, processing, and analysis. This report contains findings and recommendations derived from the discussions and presentations during the workshop.

16 TIDAL AND WAVE POWER↗

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure↗

SECARB-USA: Needs Assessment Framework for Storage Complexes (Task 2.1.b)

A team of SECARB storage experts examined 63 formations at 39 sites that were selected to represent the range of diversity of newly assessed, as well as well-advanced, storage prospects in the SECARB region. We inventoried the data needs triggered by the requirements to obtain a Class VI UIC storage permit, the data needs that results from the requirement to create geocellular fluid flow models to support that permit, and by pragmatic and best practice inputs such as public acceptance, regulatory readiness, and pore space leasing. We anchored both the needs inventory and the processes for and cost of meeting the needs with data from 9 sites in in the SECARB area that have advanced far in characterization. Results show that the total cost of characterization prior to obtaining a permit is convergent, because the permit and modeling requirements drive projects to obtain the same types of data for all cases. The high cost data that control cost are 1) drilling, coring, core-testing, logging, sampling and testing a characterization well and 2) collection of a 3-D seismic survey to map reservoir and confining system properties over the area of the plume or the area of elevated pressure. In 5 of our case study sites we determined that one or both of these costs could be avoided because the needs are met by available data.

54 ENVIRONMENTAL SCIENCES↗

Livewire: Automatic Annotations

Diogenes processes datasets to provide data quality metrics for the Livewire platform and creates standardized data dictionaries from data annotations. Diogenes needs data annotations that clearly outline thenformat and organization of the data. It also relies on the type, class, and unit of each data piece for comprehensive analysis, which it cannot determine independently. The Annotation Tool significantly reduces the time needed to create annotations for Diogenes by generating data annotations with the correct formatting and content. It also employs machine learning and hard-coded models to automatically annotate data class, quality type, and data units.

33 - ADVANCED PROPULSION SYSTEMS↗

Data Curation for Machine Learning Applied to Geothermal Power Plant Operational Data for GOOML: Geothermal Operational Optimization with Machine Learning: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach to physics-guided, data-centric machine learning. This framework has been used to develop digital twins that provide steamfield operators with operational environments to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management in real world applications. To create, test, and apply the GOOML framework, diverse time-series datasets spanning multiple years were sourced from various geothermal power plant components within several complex real-world geothermal operations. These operations are based in the United States and New Zealand and include a variety of technologies, end-uses and configurations, collectively covering nearly all relevant operating conditions for modern geothermal fields. Datasets were acquired from multiple sources to ensure that machine learning experiments generalized properly to various operating conditions. It was found that the data varied in quality, format, and completeness. To ensure consistency between the various datasets, a standardized data curation process was developed to reliably streamline data preparation. This paper will discuss best practices as learned from the GOOML data curation process which takes the following steps: 1) acquisition of large quantities of data from power plant operators, 2) digestion of data to gain an initial understanding of what is included, 3) data transformation, which includes converting the data into a standardized machine-readable format so that they can be visualized, quality checked, and cleaned, 4) quality assurance and quality control, involving identification of significant data gaps and apparent anomalies through mapping of data features to real world componentry via the GOOML historical model, followed by discussion with modelers and power plant operators to identify additional data needs and to resolve issues, 5) use in machine learning algorithms, and 6) repetition of steps one through five until all data needs are met and data are deemed suitable for producing trustworthy modeling results which may be disseminated, ideally along with the curated dataset. This iterative process is focused on improving the quality of the data rather than tuning machine learning model parameters and supports a shift towards data-centric AI as a means to improving real-world applicability of geothermal machine learning projects.

access↗

Southeast Regional CO 2 Utilization and Storage Acceleration Partnership (SECARB-USA): Needs Assessment Framework for Storage Complexes

A team of SECARB storage experts examined 63 formations at 39 sites that were selected to represent the range of diversity of newly assessed, as well as well-advanced, storage prospects in the SECARB region. We inventoried the data needs triggered by the requirements to obtain a Class VI UIC storage permit, the data needs that results from the requirement to create geocellular fluid flow models to support that permit, and by pragmatic and best practice inputs such as public acceptance, regulatory readiness, and pore space leasing. We anchored both the needs inventory and the processes for and cost of meeting the needs with data from 9 sites in in the SECARB area that have advanced far in characterization.

42 ENGINEERING↗

Modular machine learning-based elastoplasticity: Generalization in the context of limited data

The development of highly accurate constitutive models for materials that undergo path-dependent processes continues to be a complex challenge in computational solid mechanics. Challenges arise both in considering the appropriate model assumptions and from the viewpoint of data availability, verification, and validation. Recently, data-driven modeling approaches have been proposed that aim to establish stress-evolution laws that avoid user-chosen functional forms by relying on machine learning representations and algorithms. However, these approaches not only require a significant amount of data but also need data that probes the full stress space with a variety of complex loading paths. Furthermore, they rarely enforce all necessary thermodynamic principles as hard constraints. Hence, they are in particular not suitable for low-data or limited-data regimes, where the first arises from the cost of obtaining the data and the latter from the experimental limitations of obtaining labeled data, which is commonly the case in engineering applications. In this work, we discuss a hybrid framework that can work on a variable amount of data by relying on the modularity of the elastoplasticity formulation where each component of the model can be chosen to be either a classical phenomenological or a data-driven model depending on the amount of available information and the complexity of the response. The method is tested on synthetic uniaxial data coming from simulations as well as cyclic experimental data for structural materials. The discovered material models are found to not only interpolate well but also allow for accurate extrapolation in a thermodynamically consistent manner far outside the domain of the training data. This ability to extrapolate from limited data was the main reason for the early and continued success of phenomenological models and the main shortcoming in machine learning-enabled constitutive modeling approaches. Training aspects and details of the implementation of these models into Finite Element simulations are discussed and analyzed.

42 ENGINEERING↗

Advancements in Constitutive Model Calibration: Leveraging the Power of Full‐Field DIC Measurements and In Situ Load Path Selection for Reliable Parameter Inference

Accurate material characterization and model calibration are essential for computationally supported high-consequence engineering decisions. Historically, characterization and calibration methods (1) use simplified test specimen geometries and global data, (2) cannot guarantee that sufficient characterization data are collected for a specific model of interest, (3) use deterministic methods that provide best-fit parameter values with no uncertainty quantification, and (4) are sequential, inflexible, and time-consuming. This work brings together several recent advancements into an improved workflow called interlaced characterization and calibration (ICC) that advances the state-of-the-art in constitutive model calibration. The ICC paradigm (1) employs tools to efficiently use full-field data to calibrate high-fidelity material models, (2) aligns the data needed with the data collected by adopting an optimal experimental design protocol, (3) quantifies parameter uncertainty through Bayesian inference and (4) incorporates these advancements into a quasi real-time feedback loop. The ICC framework is demonstrated here on the calibration of a material model using simulated full-field data for an aluminium cruciform specimen being deformed biaxially. The cruciform is actively driven through the myopically preferred load path using Bayesian optimal experimental design, which selects load steps that yield the maximum expected information gain (EIG). Principal component analysis (PCA) is performed on the model predictions of full-field displacements, and fast surrogate models are built to approximate the input-output relationships of the expensive finite element model. Furthermore, the tools developed and demonstrated here show that high-fidelity constitutive models can be efficiently and reliably calibrated with quantified uncertainty, thus supporting credible decision-making and potentially increasing the agility of solid mechanics modelling by enabling utilization of computational simulations at earlier stages of the design cycle.

Bayesian optimal experimental design↗

Interlaced Characterization and Calibration (ICC) for Improved Computational Simulation Credibility

Accurate material characterization and model calibration are pivotal for simulations used for high-consequence engineering decisions. Current characterization and calibration methods (1) use simplified test specimen geometries and global data, (2) cannot guarantee that sufficient characterization data is collected for a specific model of interest, (3) provide only mean parameter values with no uncertainty quantification, and (4) are sequential, inflexible, and time-consuming. This work developed a new paradigm—coined Interlaced Characterization and Calibration (ICC)—which drives forward the state-of-the-art in model calibration by bringing together recent advancements into one improved workflow. The ICC paradigm (1) employs tools to efficiently use full-field data to calibrate high-fidelity material models, (2) aligns the data needed with the data collected by adopting an optimal experimental design protocol, (3) provides uncertainty metrics on the calibrated model parameters, and (4) incorporates these advances into a quasi real-time feedback loop. The ICC framework was validated synthetically with both low-fidelity and high-fidelity simulations paired with several different elastoplastic material models, and was also demonstrated experimentally with an aluminum 6061 cruciform exemplar specimen. Results showed that the ICC framework—in which Bayesian optimal experimental design actively guided the experiment— resulted in calibrations with similar or better accuracy than predetermined experiments based on subject matter expertise. Moreover, the ICC framework produced a complete model calibration— with quantified uncertainties on model parameters—in 1 week, a 5 - 10× increase in efficiency over traditional approaches. Thus, the ICC paradigm improves both the calibration process and quality, by (1) improving efficiency, which increases agility of solid mechanics modeling and enables utilization of computational simulation (CompSim) at earlier stages of the design cycle and (2) providing quantified, and in some cases reduced, parameter uncertainties, which increases confidence in model predictions and supports credible decision making.

97 MATHEMATICS AND COMPUTING↗

Design Readiness And Maturity Assessment (DRAMA) Tool for Advanced Reactors

Abstract – This research is developing a formal, repeatable method to assess the readiness and maturity of an advanced nuclear reactor design for licensing and deployment. This design readiness and maturity assessment (DRAMA) tool will be capable of determining the readiness of the design and of all parties needed to bring a particular advanced reactor design to fruition. Beneficiaries and stakeholders include the design team, research organizations (required to collect needed data and develop design tools), standards organizations, research facilities to support gathering of needed data, supply chain and construction organizations, and everyone responsible for legal and regulatory infrastructure (including defining import-export requirements). Recent experience has shown that even the most experienced engineering and construction organizations, with decades of experience in the nuclear power business, have had significant challenges in bringing designs to completion, licensing the designs, and constructing new plants. For new entries into the field, simply understanding the unique environmental and regulatory requirements have been daunting. A significant part of the challenge is that new entries into the market do not know what they don't know. This tool will provide design teams a better understanding of their readiness to proceed to licensing (and other steps in the process), while at the same time providing a valuable metric for other interested organizations, such as funding agencies, national regulators, and international markets. The DRAMA tool will provide an assessment of the likelihood of successfully completing licensing and deployment and will also create the capability to assess the ability of regulatory infrastructure and the supply chain to support the deployment of any design or class of designs. It will also help prioritize research and policy efforts to improve the likelihood of deployment of the next generation of advanced reactors.

Arndt, Steven↗

First Report of the Nuclear Data Subcommittee of the Nuclear Science Advisory Committee

Accurate, reliable nuclear data is essential for the success of Federal missions such as nonproliferation, nuclear forensics, homeland security, national defense, space exploration, clean energy generation, and scientific research. Data access is also key to innovative commercial developments such as new medicines, automated industrial controls, energy exploration, energy security, nuclear reactor design, and isotope production. The United States Nuclear Data Program (USNDP) is the domestic custodian of nuclear data. In its April 2022 meeting, the DOE/NSF Nuclear Science Advisory Committee was charged with preparing two reports on nuclear data. In this first report, we review recent accomplishments of the USNDP and discuss complementary and collaborative international efforts. Detailed descriptions of nuclear data needs for basic science, nonproliferation, national security, nuclear energy together with medical and space applications are also presented. Lastly, a set of specific cross-cutting nuclear data needs with relevance for multiple applications areas are also identified for further discussion in a follow-on report planned for release at the end of January 2023.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Evaluations for medium- and high-mass nuclei for FUSION applications

There is a renewed attention to nuclear fusion as a commercial source of carbon-free energy, however there are many scientific needs that must be addressed to enable the future success of fusion as an economical energy option. Among these is the proper description of the impact of radiation produced in the fusion vessel chamber and all other components of the reactor. In this work we will focus on the nuclear data needs to describe the interaction between primary and secondary neutron radiation and the medium- and high-mass nuclei commonly present in structural (such as stainless steel) and superconducting (e.g., electromagnets) materials.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Corrective Action Decision Document/Closure Report for Corrective Action Unit 542: Disposal Holes, Nevada Test Site, Nevada, Revision 0, with ROTC 1-2

The purpose of this Corrective Action Decision Document/Closure Report is to provide justification and documentation supporting the recommendation for closure of CAU 542 with no further corrective action. To achieve this, corrective action investigation (CAI) activities were performed from July 19 through August 25, 2006, as set forth in the Corrective Action Investigation Plan for Corrective Action Unit 542: Disposal Holes. The purpose of the CAI was to fulfill the following data needs as defined during the data quality objective (DQO) process: Determine whether contaminants of concern (COCs) are present. If COCs are present, determine the nature and extent. Provide sufficient information and data to complete appropriate corrective actions.

54 ENVIRONMENTAL SCIENCES↗

An Open-Source Framework for Characterizing Urban Energy Models: Integrating Top-Down and Bottom-Up Methods to Predict Residential Buildings Characteristics: Preprint

Bottom-up urban energy models are crucial for understanding current energy use patterns and informing design strategies. However, accurately characterizing these models to represent different communities remains a challenge due to the extensive data needed for simulating existing energy use behavior. This data includes information related to human activities and building characteristics, all of which correlate with socioeconomic factors. To overcome this challenge, we developed an automated framework that utilizes both top-down and bottom-up data, to predict unknown building and occupant characteristics that are needed for more accurate and equitable modeling and analytics. Our framework, integrated into the URBANopt district energy modeling platform, uses statistical data models from ResStock. URBANopt models co-located buildings and neighborhoods. At this scale there are data gaps in building characteristic data, such as materials, insulation, occupancy, income, and energy usage of the buildings. To address this data gap, we use ResStock data, representative at the census tract scale, and develop machine-learning and deeplearning techniques to disaggregate it to individual buildings. By mapping unique occupant, building and economic properties to URBANopt energy models, we gain detailed insights into the variability of building energy use across different neighborhoods. This insight helps deploy technologies for co-located buildings and supports targeted upgrades for communities with unique economic and demographic characteristics, ensuring energy equity. Accurate characterization of energy models allows us to develop equitable strategies tailored to diverse neighborhoods, whether underserved or affluent. Our automated framework streamlines energy modeling and provides a reliable tool for building energy characterization.

ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATION↗

Workshop Summary: Bridging the Gap Between Atmospheric Science and Grid Integration

The need for dedicated, accurate, expertly curated weather data is increasingly important as the share of variable renewable energy increases on the power system. Projections for futures with very high (50+% annual energy) shares of variable generation require ongoing assessment of data requirements from industry stakeholders in their power system operation and planning contexts. In March 2024, NREL organized a workshop entitled "Bridging the Gap Between Atmospheric Science and Grid Integration Workshop", which brought atmospheric scientists and power system experts together to refine the requirements of atmospheric datasets for grid integration, and to describe a holistic approach to creating new and regularly updated national scale wind datasets for power system planning and operations. The results of this workshop are being used to inform the near-term development and a longer-term strategy for DOE to produce relevant wind resource datasets and inform wider use of wind/solar/load data sets in power system planning. This presentation provides an overview of a preworkshop survey, an assessment of current state of the art of national-scale datasets for wind resource assessment and grid integration, insights on appropriate uses of the WTK-LED, power system perspectives on data needs, as well as recommended next steps as discussed in the workshop and how these steps support longer-term strategies.

17 WIND ENERGY↗

Oscilloscope Data Push Program

This paper details the development of a Python program designed to automate the data acquisition and conversion for an oscilloscope for the purposes of a one-off/temporary data acquisition system for users that readily need data, and do not have the option of obtaining a Data Acquisition (DAQ) solution. Creating DAQ systems for analyzing a system requires expensive electronics and a dedicated team of engineers for support. Traditionally, manual data collection and processing are time consuming and prone to error. By automating these processes, the cost, efficiency and accuracy of data handling are improved upon. This project involves the creation of a program that interacts with the oscilloscope. During this interaction, there are various functions being performed such as the acquisition of waveform data via floating points, generating plots with the acquired wave points, and storing of floating points in a CSV file format for future reference and plotting purposes. While the initial aim of the project included continuous logging to a cloud database, this was deferred due to time constraints. The results portrayed an almost-instant rate of data collection with a buffer time, showcasing the potential for further integration and real-time data processing.

Osei-Tutu, Jason↗

Jobs, jobs, jobs: what’s an analyst to do?

Analysts and economists often face the task of using employment metrics to characterize industries of interest. Some key challenges can be understanding where to find employment metrics, the differences in various employment metrics, and when each metric should be used. This article analyzes a variety of publicly available employment data for the United States and compares these data. A detailed description of the intricacies of each data source is provided, which covers factors such as regionality, industry breakout, periodicity, and the types of jobs included. This article provides several case study examples, using the oil and gas extraction, coal mining, and chemical manufacturing sectors to portray challenges data users may face when developing employment estimates that suit their needs. Data users should be aware of a variety of data sources to understand alternative analysis options when data limitations are present and to determine which data source best meets their needs. Instances may occur in which information from one dataset may be used to help impute missing values.

99 GENERAL AND MISCELLANEOUS↗