Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “standardized data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

2024 Standard Scenarios: A U.S. Electricity Sector Outlook

This data corresponds to the 2024 Standard Scenarios report, which contains a suite of forward-looking scenarios of the possible evolution of the U.S. electricity sector through 2050. These files contain modeled projections of the future. Although we strive to capture relevant phenomena as comprehensively as possible, the models used to create this data are unavoidably imperfect, and the future is highly uncertain. Consequentially, this data should not be the sole basis for making decisions. In addition to drawing from multiple scenarios within this set, we encourage analysts to also draw on projections from other sources, to benefit from diverse analytical frameworks and perspectives when forming their conclusions about the future of the power sector. For further discussions about the limitations of the models underlying this data, see section 1.4 of the "ReEDS Documentation" linked below. For scenario descriptions, input assumptions, and metric definitions for the data in these files, see the "2024 Standard Scenarios Report" linked below.

2050↗

Ontologizing health systems data at scale: making translational discovery a reality

Common data models solve many challenges of standardizing electronic health record (EHR) data but are unable to semantically integrate all of the resources needed for deep phenotyping. Open Biological and Biomedical Ontology (OBO) Foundry ontologies provide computable representations of biological knowledge and enable the integration of heterogeneous data. However, mapping EHR data to OBO ontologies requires significant manual curation and domain expertise. We introduce OMOP2OBO, an algorithm for mapping Observational Medical Outcomes Partnership (OMOP) vocabularies to OBO ontologies. Using OMOP2OBO, we produced mappings for 92,367 conditions, 8611 drug ingredients, and 10,673 measurement results, which covered 68–99% of concepts used in clinical practice when examined across 24 hospitals. When used to phenotype rare disease patients, the mappings helped systematically identify undiagnosed patients who might benefit from genetic testing. By aligning OMOP vocabularies to OBO ontologies our algorithm presents new opportunities to advance EHR-based deep phenotyping.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A Python Instrument Control and Data Acquisition Suite for Reproducible Research

Tools that standardize and automate experimental data collection are needed for greater confidence in research results. The National Synchrotron Light Source-II (NSLS-II) has generated an open-source Python data acquisition, management, and analysis software suite that automates X-ray experiments and collects an experimental record that facilitates complete reproducibility. Here, we show that the NSLS-II tools are not only useful for X-ray science at large-scale facilities by presenting an add-on package that adapts these tools for use in a small laboratory with common physics and electrical engineering instruments. The composite software suite eases and automates the execution of experiments, records extensive metadata, stores data in portable containers, and speeds up the analysis through tools for comprehensive searches. In total, this software suite increases the reproducibility of laboratory experiments. We demonstrate the software via the evaluation of two lock-in amplifiers-the miniature ADA2200 and the ubiquitous Stanford Research Systems (SRS) SR810. The frequency resolution, signal-to-noise ratio, and dynamic reserve of the lock-in amplifiers are measured and presented. The usage of the software suite is described throughout these measurements so that the reader can implement the tools in their lab.

97 MATHEMATICS AND COMPUTING↗

The Linear Point Standard Ruler with DESI DR1 and DR2 Data

The linear point, a purely geometric feature in the monopole of the two-point correlation function, has been proposed as an alternative standard ruler. Compared to the peak in the correlation function, it is more robust to late-time nonlinear effects at the percent level. In light of improved simulations and high quality data, we revisit the robustness of the linear point and use it as an alternative to template-based fitting approaches typically used in BAO analyses. We present the linear point measurements on galaxy samples from the first and second data releases (DR1 and DR2) of the DESI survey. We convert the linear point into a dimensionless parameter $α_{iso,LP}$, defined as the ratio of the linear point in the fiducial cosmology and the observed value, analogous to the isotropic BAO scaling parameter $α_{iso}$ used in previous BAO measurements. Using the 2nd generation of AbacusSummit mock catalogs, we find that linear point measurements are more precise when calculated in the post-reconstruction regime with 15-60% smaller uncertainties than those pre-reconstruction. We find a systematic shift in the linear point measurements compared against the isotropic BAO measurements in mocks; we attribute this to the isotropic damping parameter responsible for smearing the linear point in the nonlinear regime. We propose a sample-dependent correction that mitigates the impact of late-time nonlinear effects. While this introduces a cosmology dependence in an otherwise model-independent measurement, this is necessary given the sub-percent precision dictated by current cosmological surveys. Comparing $α_{iso,LP}$ with isotropic BAO measurements made on the DESI DR1 and DR2 galaxy samples, we find excellent agreement after applying this correction, particularly post-reconstruction. We discuss future scope regarding cosmological inference with linear point measurements.

Uberoi, N. [Yale U.] (ORCID:0000000275179629)↗

HPC ODA Commons [SWR-26-003]

HPC ODA Commons is a community-driven platform for standardizing HPC operational data analytics. HPC sites generate enormous volumes of operational data - scheduler logs, accounting records, monitoring streams - but turning that data into actionable insight is needlessly hard. Each site builds bespoke parsers, schemas, and evaluation pipelines. Results can't be compared across institutions. Promising analytics ideas stay siloed because there's no shared language for describing the data, the experiments, or the outcomes. HPC ODA Commons fixes this by establishing community-governed contracts - versioned schemas, canonical artifacts, and benchmark recipes - that make ODA workflows discoverable, reproducible, and comparable. It pairs these standards with a practical, CLI-first toolkit that lets operators and researchers go from raw logs to standardized results without sending data off-cluster.

Menear, Kevin [National Laboratory of the Rockies ↗

Instrumentation and methods for efficient time-resolved X-ray crystallography of biomolecular systems with sub-10 ms time resolution

Time-resolved X-ray crystallography has great promise to illuminate structure–function relations and key steps of enzymatic reactions with atomic resolution. The dominant methods for chemically-initiated reactions require complex instrumentation at the X-ray beamline, significant effort to operate and maintain this instrumentation, and enormous numbers (∼10 5 –10 9 ) of crystals per time point. We describe instrumentation and methods that enable high-throughput time-resolved study of biomolecular systems using standard crystallography sample supports and mail-in X-ray data collection at standard high-throughput cryocrystallography synchrotron beamlines. The instrumentation allows rapid reaction initiation by mixing of crystals and substrate/ligand solution, rapid capture of structural states via thermal quenching with no pre-cooling perturbations, and yields time resolutions in the single-millisecond range, comparable to the best achieved by any non-photo-initiated method in both crystallography and cryo-electron microscopy. Our approach to reaction initiation has the advantages of simplicity, robustness, low cost, adaptability to diverse ligand solutions and small minimum volume requirements, making it well suited to routine laboratory use and to high-throughput screening. We report the detailed characterization of instrument performance, present structures of binding of N -acetylglucosamine to lysozyme at time points from 8 ms to 2 s determined using only one crystal per time point, and discuss additional improvements that will push time resolution toward 1 ms.

Indergaard, John A. (ORCID:000000022367699X)↗

labquake_future_prediction

The labquake_future_prediction code is a collection of python modules and scripts that serves as supporting information for the article “Predicting future laboratory fault friction through deep learning” for publication in the journal of “Geophysical Research Letters”. It is designed to predict laboratory fault slips in the immediate future by scanning continuous acoustic emission (AE) waveforms recorded in laboratory biaxial shear experiments. The predictions are made with a deep learning model based on convolutional encoder-decoder (CED) models and the Transformer model primarily developed for Natural Language Processing (NLP). The deep learning model is trained with the tensorflow package using publicly available laboratory data sets in standard binary file format in numpy. The utility functions for reading data files, configuring model hyperparameters, constructing the CED and Transformer models, training and testing of the models are defined in python module files. The workflow of training the models for labquake future predictions and the multiple GPU’s rapid model hyperparameter optimization as described in the journal article, are demonstrated in accompanying python script files and Jupyter notebooks.

Wang, Kun↗

Measurement Uncertainty in One-Of-A-Kind Event Data Analysis

A golden standard in science is to repeat an experiment a statistically significant number of times, recording data using the same set of detectors and the same data analysis methodology. In such case experimental error includes both the range of true values generated by repetitions of the experiment, and measurement uncertainty caused by the detector. They are independent. It is a huge and too frequently used simplification, to assume that one can measure multiple repetitions of an identical experiment, resulting in identical true experimental value. Repetitions, as similar is it is experimentally achievable, have unavoidable built-in differences resulting in a range of the true values rather than in a single value. When modern, very sensitive and well calibrated measurement systems are used, this range is not negligible, and sometimes dominates over the measurement uncertainty. Range of true values depends on built-in differences in physics of the experiment. Stochastic physical processes result typically in a broader range of true values than non-stochastic processes do. Measurement uncertainty depends on a measurement method (properties of the detector not of the experiment). Modern measurement methods, including digital ones, frequently make the measurement uncertainty very small. When data from one–of –a kind experiment are analyzed, only the measurement uncertainty is reported. It provides no information about the range of true experimental values, neither about reliability of a reported data point. Reliability of a data point is in general independent from its measurement uncertainty. However, in practice reliable measurement methods frequently have high measurement uncertainty, while low reliability methods are applied to limit measurement uncertainty. Comparison of reliable data with high measurement uncertainty to not so reliable data measured with low uncertainty is discussed – in different scenarios different data analysis methods are applicable. Methods for data analysis from an experiment repeated statistically significant number of times are very well developed. They do not require a detailed expertise in physics of an experiment, nor in the properties of the measurement system used, and meaning of the reported uncertainty is well understood in any scientific community. It all changes when data from one-of-a-kind experiment is analyzed. Analyst’s expertise is required both in the physics of the experiment and in all aspects of the measurement system, all possible malfunctions. Data users must remember that only measurement uncertainty is reported from any one-of-a-kind experiment. Theory with simulations may provide estimation of expected built-in differences in the experiment, and by this of expected range of true values for a given experiment; yet measurement uncertainty can never be used in place of the range of true experimental values.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Data Deposition Platform for Sharing Nuclear Magnetic Resonance Data

Nuclear magnetic resonance (NMR) data are rarely deposited in open databases, leading to loss of critical scientific knowledge. Existing data reporting methods (images, tables, lists of values) contain less information than raw data, and are poorly standardized. Together, these issues limit FAIR (findable, accessible, interoperable, reusable) access to these data, which in turn creates barriers for compound dereplication and the development of new data-driven discovery tools. Existing NMR databases are either not designed for natural products data, or employ complex deposition interfaces that disincentivize deposition. Journals, including the Journal of Natural Products (JNP), are now requiring data submission as part of the publication process, creating the need for a streamlined, user-friendly mechanism to deposit and distribute NMR data. Recently, our team reported the development of the Natural Products Magnetic Resonance Database (NP-MRD; www.np-mrd.org). Here in this paper we present a new data deposition platform for the NP-MRD project that is designed to enable users to deposit NMR data for published or submitted manuscripts in under five minutes. This platform includes a suite of automated data extraction and standardization tools, together with a simple-to-use web-based interface and detailed error reporting to simplify the data deposition process and is available at www.np-mrd.org/submissions.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A System for Standardizing and Combining U.S. Environmental Protection Agency Emissions and Waste Inventory Data

The U.S. Environmental Protection Agency (USEPA) provides databases that agglomerate data provided by companies or states reporting emissions, releases, wastes generated, and other activities to meet statutory requirements. These databases, often referred to as inventories, can be used for a wide variety of environmental reporting and modeling purposes to characterize conditions in the United States. Yet, users are often challenged to find, retrieve, and interpret these data due to the unique schemes employed for data management, which could result in erroneous estimations or double-counting of emissions. To address these challenges, a system called Standardized Emission and Waste Inventories (StEWI) has been created. The system consists of four python modules that provide rapid access to USEPA inventory data in standard formats and permit filtering and combination of these inventory data. When accessed through StEWI, reported emissions of carbon dioxide to air and ammonia to water are reduced approximately two- and four-fold, respectively, to avoid duplicate reporting. StEWI will greatly facilitate the use of USEPA inventory data in chemical release and exposure modeling and life cycle assessment tools, among other things. To date, StEWI has been used to build the recent USEEIO model and the baseline electricity life cycle inventory database for the Federal LCA Commons.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Soft costs and EVSE – Knowledge gaps as a barrier to successful projects

There has been a recent push to increase access to electric vehicle (EV) charging infrastructure. The National Electric Vehicle Infrastructure (NEVI) program, part of the Bipartisan Infrastructure Law (BIL) has made significant funding available for major charging infrastructure projects along state thruways, and many state and local incentives exist for EV owners to install chargers in their homes. However, deployment of these chargers has not kept up with demand, primarily due to issues in project planning, permitting processes, and unforeseen delays. This paper serves as a review of the current understanding of these and other non-hardware costs in EV charging infrastructure projects (collectively known as “soft costs”). We found that soft costs in EV charging infrastructure projects are not well understood. Specifically, there is little agreement on how soft costs should be categorized and tracked, and less agreement still on best practices for controlling these costs and lowering barriers to infrastructure deployment. A broader review of EV charging infrastructure cost analyses shows that these costs can have significant impacts on project outcomes. EV charging infrastructure projects may be able to examine the success of the solar industry in lowering soft costs, and a similar effort may lower project costs significantly. Further work on standardizing and collecting data on EV charging infrastructure costs is required to begin addressing and controlling these costs.

32 - ENERGY CONSERVATION, CONSUMPTION, AND UTILIZA↗

Schema Elements for Granta Annual Report: FY2024

Granta: Materials Intelligence (Granta: MI) is a commercial database software distributed by Ansys, Inc. that is utilized by the Nuclear Security Enterprise (NSE) to organize and store relevant materials data. Lack of standard and well-documented database schema is the primary obstacle to an NSE materials data management solution, so the objective of this project is to create and document such a schema. In FY21, an approach for designing, documenting, and managing a standard database schema was described based on the creation of schema elements (collections of attributes used to describe particular aspects of the data) to be used as building blocks for creating various database tables without duplication. In FY22, these methods were applied through a multi-site collaboration to create and document the schema elements necessary to build a thermogravimetric analysis (TGA) testing table. In FY23 the schema was expanded to include elements for a differential scanning calorimetry (DSC) table, along with schema for supporting metadata tables including Instruments, Projects, Documents, and Testing Series. In FY24 the following progress was made, again through multi-site collaboration: • The existing schema elements were modified to accommodate thermomechanical analysis (TMA) data, and a table, Test Data: TMA, was created for managing TMA data. • The elements necessary for the following additive manufacturing (AM) data tables (directed at data specific to selective laser sintering AM technology) were created: • AM Builds • AM Processes • AM Part Designs • Built AM Parts • AM Feedstock Materials • AM Feedstock Material Batches • The elements necessary for creating a Calibrated Material Models table were created, and the Calibrated Material Models table was created. In FY25 the existing schema will be deployed on the production enterprise Granta instance on the enterprise secure network. Schema elements will be appended, and new elements created as necessary, to allow the creation of tables specifically to support materials testing, AM process development, and design and analysis for modernization programs.

36 MATERIALS SCIENCE↗

Creation of a Weather Drivers Test Suite for Inclusion in ASHRAE Standard 140

Weather conditions are an important boundary condition for building performance simulation (BPS) calculations. For existing test cases in ASHRAE Standard 140 "Method of Test for Evaluating Building Performance Simulation Software" (ANSI/ASHRAE 2020), it was assumed that the software being tested could adequately read and interpret the weather data in the provided standard weather files. As differences between the programs have been reduced and as more programs have shifted to sub-hourly time steps this assumption has become more stretched. To address these concerns a new test suite testing a program's ability to read and interpret the data from a standard weather file was developed. The purpose of the test suite is to test the use of the typical data used from standard weather files.

54 ENVIRONMENTAL SCIENCES↗

Modeling, Validation, and Control of the IEA‐15 MW Reference Wind Turbine and VolturnUS‐S Platform

This paper presents the acausal modeling, validation, and control of floating offshore wind turbines (FOWTs). The model simulates the IEA‐15 MW reference turbine and the semi‐submersible VolturnUS‐S platform utilizing a Control‐oriented, Reconfigurable, and Acausal Floating Turbine Simulator (CRAFTS), which integrates the key coupled aero‐hydro‐elasto‐servo dynamics and is being developed by authors at the University of Central Florida. Verification and validation are conducted using numerical data from the industry‐standard simulation platform OpenFAST and experimental data from the Floating Offshore‐wind and Controls Advanced Laboratory (FOCAL) project, in which the authors were involved. Numerical results demonstrate the model's ability to qualitatively capture loads and responses across various load cases, highlighting the impact of the control system under different wind and wave conditions and opening new opportunities for optimizing FOWT designs. This paper provides wind turbine researchers with valuable insights into system characteristics, system frequencies, damping effects, and internal reaction forces, serving as a reference for future studies in FOWT modeling and control.

17 WIND ENERGY↗

Assessing the Consistency of Estimated Ground Cover Fractions between the BLM AIM Method and Optical Remote Sensing Method

Since 2012, Argonne National Laboratory (Argonne) has supported the Bureau of Land Management (BLM) in developing remote sensing methodologies for long-term environmental monitoring of Palo Verde Mesa in eastern Riverside County, California, including methods for: detailed mapping of ephemeral streams, estimating fractional cover of desert-land surface components (e.g., trees, shrubs, litters, and bare ground), evaluating erosion risk or land stability, and characterizing vegetation alliances using spatial structure and geostatistical approaches. These studies showed the promise of remote sensing for monitoring changes in desert landscapes by providing information that would be difficult to obtain through field surveys. During this time the BLM has also worked to establish long-term monitoring protocols and compiled field-observation data collected using standardized protocols from the Assessment, Inventory, and Monitoring (AIM) strategy. The AIM data can be compared to data derived from remote sensing methods to evaluate their relative operational utility in monitoring landscape change. If ground cover estimated using remotely sensed imagery is comparable to AIM ground cover estimates, then remote sensing can be used to monitor whether any land cover change in desert landscapes may be related to solar energy development. Therefore, the goal of this study was to determine the consistency in ground cover estimates between AIM data and those derived from publicly- available remotely sensed imagery, such as that available through the U.S. Department of Agriculture, National Agricultural Imagery Program (NAIP), to examine feasibility of a remote sensing method for complementing AIM monitoring. Based on the image analysis in this study, we also provide recommendations for how small unmanned aerial system (sUAS) data may be used to complement BLM’s AIM data and NAIP imagery for future vegetation monitoring. The ground cover types we originally planned to investigate were trees, shrubs, and bare ground. However, the small sample size and a limited range of cover fraction of trees and shrubs in the AIM dataset (e.g., 33 samples with a maximum shrub cover of 18%, 14 samples with a maximum tree cover of 17%) did not allow for performing a meaningful evaluation for the remote sensing approach. Therefore, we conducted the study focusing on bare ground, foliar, and rock cover, all of which are indicators reported in the AIM remote sensing dataset.

47 OTHER INSTRUMENTATION↗

The Neurodata Without Borders ecosystem for neurophysiological data science

The neurophysiology of cells and tissues are monitored electrophysiologically and optically in diverse experiments and species, ranging from flies to humans. Understanding the brain requires integration of data across this diversity, and thus these data must be findable, accessible, interoperable, and reusable (FAIR). This requires a standard language for data and metadata that can coevolve with neuroscience. We describe design and implementation principles for a language for neurophysiology data. Our open-source software (Neurodata Without Borders, NWB) defines and modularizes the interdependent, yet separable, components of a data language. We demonstrate NWB’s impact through unified description of neurophysiology data across diverse modalities and species. NWB exists in an ecosystem, which includes data management, analysis, visualization, and archive tools. Thus, the NWB data language enables reproduction, interchange, and reuse of diverse neurophysiology data. More broadly, the design principles of NWB are generally applicable to enhance discovery across biology through data FAIRness.

59 BASIC BIOLOGICAL SCIENCES↗