Engineering PapersSearch

SEARCH · Engineering Papers

Results for “analysis ready data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

End-Use Savings Shapes Upgrade Package Documentation: LED Lighting, HP-RTU and ASHP-Boiler

Building on the successfully completed effort to calibrate and validate the U.S. Department of Energy's ResStock and ComStock models over the past 3 years, the objective of this work is to produce national data sets that empower analysts working for federal, state, utility, city, and manufacturer stakeholders to answer a broad range of analysis questions. The goal of this work is to develop energy efficiency, electrification, and demand flexibility end-use load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) upgrade measures, or upgrades. "Measures" refers to energy efficiency variables that can be applied to buildings during modeling. An end-use savings shape is the difference in energy consumption between a baseline building and a building with an energy efficiency, electrification, or demand flexibility upgrade applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step. ComStock is a highly granular, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the commercial building stock across the United States. The baseline model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology and results of the baseline model are discussed in the final technical report of the End-Use Load Profiles project. This documentation focuses on an upgrade package of three end-use savings shapes upgrades - Light Emitting Diode (LED) Lighting, Heat Pump Rooftop Unit (RTU) (HP-RTU), and Air-Source Heat Pump (ASHP) Boiler, which we will refer to collectively as the "Interior Lighting and Heat Pump" package. More details on the individual upgrades can be found on the ComStock Measures Documentation page. An upgrade package applies two or more EUSS upgrades to a single building model simulation. Since ComStock is a bottom-up physics-based model, an upgrade package will go beyond aggregating or summing the individual upgrade results and produce novel results by simulating interactions between the upgrades. For example, pairing an envelope upgrade with an electrification upgrade would likely result in higher savings results than the sum of these upgrades individually, and the size of the heating, ventilating, and air conditioning (HVAC) equipment may be reduced if the envelope upgrade reduces the loads significantly.

29 ENERGY PLANNING, POLICY, AND ECONOMY

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

A Modelica Implementation of an Organic Rankine Cycle

Organic Rankine cycle (ORC) systems generate power from low-grade heat sources, such as geothermal sources and industrial waste heat. A key feature is that a working fluid is selected to match the temperature of the source. With the vast pool of candidate working fluids comes the challenge of developing a large number of robust thermodynamic media models. We implemented a subcritical ORC model in Modelica that uses working fluid data records and interpolation schemes in lieu of thermodynamic medium evaluation for energy recovery estimation. This is a component model that can be integrated into a larger energy system model. It does not require detailed thermodynamic, heat transfer, or machine analysis. Our ORC model fills a gap where working fluids are ready to choose or easy to add, and at the same time can be integrated into an energy system.

29 ENERGY PLANNING, POLICY, AND ECONOMY

End-Use Savings Shapes Upgrade Package Documentation: Wall and Roof Insulation, New Windows, LED Lighting, HP-RTU and ASHP-Boiler

Building on the successfully completed effort to calibrate and validate the U.S. Department of Energy’s ResStock™ and ComStock™ models over the past 3 years, the objective of this work is to produce national data sets that empower analysts working for federal, state, utility, city, and manufacturer stakeholders to answer a broad range of analysis questions. The goal of this work is to develop energy efficiency, electrification, and demand flexibility enduse load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) upgrade measures, or upgrades. “Measures” refers to energy efficiency variables that can be applied to buildings during modeling. An end-use savings shape is the difference in energy consumption between a baseline building and a building with an energy efficiency, electrification, or demand flexibility upgrade applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Validation of the DESI DR2 measurements of baryon acoustic oscillations from galaxies and quasars

The Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) galaxy and quasar clustering data represents a significant expansion of data from Data Release 1 (DR1), providing improved statistical precision in baryon acoustic oscillation (BAO) constraints across multiple tracers, including bright galaxies, luminous red galaxies, emission line galaxies, and quasars. In this paper, we validate the BAO analysis of DR2. We present the results of robustness tests on the blinded DR2 data and, after unblinding, consistency checks on the unblinded DR2 data. All results are compared with those obtained from a suite of mock catalogs that replicate the selection and clustering properties of the DR2 sample. We confirm the consistency of DR2 BAO measurements with DR1 while achieving a reduction in statistical uncertainties due to the increased survey volume and completeness. The combined BAO precision, including both statistical and systematic errors, improves from ∼0.52% in DR1 to 0.30% in DR2—a factor of 1.7 gain. We assess the impact of analysis choices, including different data vectors (correlation function vs power spectrum), modeling approaches and systematics treatments, and an assumption of the Gaussian likelihood, finding that our BAO constraints are stable across these variations and assumptions with a few minor refinements to the baseline setup of the DR1 BAO analysis. We summarize a series of pre-unblinding tests that confirmed the readiness of our analysis pipeline, the final systematic errors, and the DR2 BAO analysis baseline. The successful completion of these tests led to the unblinding of the DR2 BAO measurements, ultimately leading to the DESI DR2 cosmological analysis, with their implications for the expansion history of the Universe and the nature of dark energy presented in the DESI key paper (companion paper).

79 ASTRONOMY AND ASTROPHYSICS

NanoPSD: A software for automatic detection of Nano-Particle Shape Distribution in electron microscopy images

Accurate quantification of the size and morphology of nanoparticles from electron microscopy (EM) images is essential to understand growth mechanisms, surface reactivity, and functional behavior in nanoscale materials. Manual analysis remains slow, subjective, and difficult to reproduce in large datasets. We introduce NanoPSD (Nano-Particle Shape Distribution), an open-source and fully automated framework for quantitative particle detection and morphology analysis from EM images. NanoPSD integrates adaptive contrast enhancement, polarity-agnostic scale-bar detection, Optical Character Recognition (OCR)-based calibration, and classical segmentation via Otsu thresholding with morphological refinement. Particle contours are used to extract geometric descriptors, including equivalent circular diameter, aspect ratio, circularity, and solidity, enabling automated classification into spherical, rod-like, and aggregate morphologies. The framework supports both single-image and batch processing, generating publication-quality visualizations, LaTeX-ready tables, and structured comma-separated values (CSV) datasets. As a demonstration, we applied NanoPSD to plasma-synthesized nanoparticle samples diagnosed via transmission electron microscopy (TEM). The code produced statistically robust size and morphology distributions spanning a few to tens of nanometers with minimal user supervision. The pipeline demonstrates high reproducibility and scalability, processing large image collections with consistent calibration and output formatting. Its modular design enables seamless integration of future deep-learning-based segmentation models, providing a pathway toward intelligent, data-driven electron microscopy analysis.

36 MATERIALS SCIENCE

SoK: What does it Mean to Benchmark Database Forensics?

Relational Database Management Systems are the backbone of modern enterprises and public-sector services, and are thus frequent targets of security incidents, insider threats, and thorough regulatory audits. Consequently, databases have become key sources of digital evidence, requiring investigators to reconstruct past activity from audit logs, transaction logs, and backups. Although benchmarking frameworks such as those developed by the Transaction Processing Performance Council (TPC) are widely used to evaluate database performance, they do not capture forensic requirements such as evidentiary completeness, tamper-evidence, chain of custody, or regulatory compliance under GDPR and CCPA. This survey examines the emerging domain of forensic database benchmarking. We gathered prior research on database forensics, secure logging, and tamper-evident data structures; we analyze modern forensic-ready features in commercial and open-source systems (SQL Server Ledger, Oracle Blockchain Tables, PostgreSQL pgAudit, Db2 Audit, Aurora Database Activity Streams, Oracle Real Application Security and IBM Guardium) and assess why existing benchmarks are insufficient. We propose forensic workloads, metrics, and methodologies that incorporate adversarial stressors, deleted-record recovery, and backup analysis. We also identify open research problems and call for a community-driven forensic benchmark suite. The result is an idea for evaluating not only database performance but also forensic soundness, bridging the gap between system engineering, compliance, and digital investigations.

Lenard, Ben

SERA: A Hydrogen Infrastructure Capacity Expansion Model

The Scenario Evaluation and Regionalization Analysis (SERA) model is an infrastructure planning optimization model that can guide hydrogen production, delivery, and end-use investment decisions and accelerate the adoption of low-cost hydrogen at scale, whether for fuel cell electric vehicles or non-transportation applications. In this talk, we will review the SERA model objective function as well as the data inputs and outputs. We will also look at a SERA case study identifying potential dispensed costs of hydrogen along major refueling corridors throughout the United States. In addition to the SERA model, Justin will also discuss his recent work for the Office of Manufacturing and Energy Supply Chains on electrolyzer supply chain readiness, and his work for the Hydrogen Fuel Cell Technologies Office and Environmental Protection Agency on the levelized cost of dispensed hydrogen for heavy-duty trucking.

30 DIRECT ENERGY CONVERSION

Analysis of Hydrogen Supply Chain Readiness in Selected Indo-Pacific Countries

This report explores the low-carbon hydrogen production and procurement targets of Indo-Pacific Economic Framework for Prosperity (IPEF) countries, highlighting key cost and infrastructure considerations. By 2030, production costs are targeted between US$\$$1.4-4.8/kg-H2, with procurement prices at US$\$$2.3/kg-H2, narrowing further by 2050. Midstream costs, including storage and transportation, are critical to meeting these targets and enabling hydrogen trade. The report discusses challenges such as infrastructure readiness, regulatory frameworks, and the need for improved data reporting. It emphasizes the importance of tax incentives, research investments, and port readiness to support the growth of low-carbon hydrogen markets in the region.

08 HYDROGEN

PSTN-019: The LSST Science Pipelines Software: Optical Survey Pipeline Reduction and Analysis Environment

The NSF-DOE Vera C. Rubin Observatory is executing the Legacy Survey of Space and Time (LSST) as its prime mission, producing a series of data releases over the ten-year survey. The LSST Science Pipelines Software will be used to create these data releases and to perform the nightly prompt processing and alert production. This paper provides an overview of the LSST Science Pipelines Software, describing the components and their integration into pipelines that generate science-ready data products.

79 ASTRONOMY AND ASTROPHYSICS

High-throughput single-cell transcriptomics of bacteria using combinatorial barcoding

Microbial split-pool ligation transcriptomics (microSPLiT) is a high-throughput single-cell RNA sequencing method for bacteria. With four combinatorial barcoding rounds, microSPLiT can profile transcriptional states in hundreds of thousands of Gram-negative and Gram-positive bacteria in a single experiment without specialized equipment. As bacterial samples are fixed and permeabilized before barcoding, they can be collected and stored ahead of time. During the first barcoding round, the fixed and permeabilized bacteria are distributed into a 96-well plate, where their transcripts are reverse transcribed into cDNA and labeled with the first well-specific barcode inside the cells. The cells are mixed and redistributed two more times into new 96-well plates, where the second and third barcodes are appended to the cDNA via in-cell ligation reactions. Finally, the cells are mixed and divided into aliquot sub-libraries, which can be stored until future use or prepared for sequencing with the addition of a fourth barcode. It takes 4 days to generate sequencing-ready libraries, including 1 day for collection and overnight fixation of samples. Here, the standard plate setup enables single-cell transcriptional profiling of up to 1 million bacterial cells and up to 96 samples in a single barcoding experiment, with the possibility of expansion by adding barcoding rounds. The protocol requires experience in basic molecular biology techniques, handling of bacterial samples and preparation of DNA libraries for next-generation sequencing. It can be performed by experienced undergraduate or graduate students. Data analysis requires access to computing resources, familiarity with Unix command line and basic experience with Python or R.

59 BASIC BIOLOGICAL SCIENCES

Hydrogen Infrastructure Analysis for the Port Applications [Slides]

The International Maritime Organization has committed to 50% reduction in GHG emissions by 2050 worldwide as of 2023. This analysis includes performing an inventory and modeling efforts to understand the energy, equipment and cost requirements to support decarbonization of cargo handling and shore power at U.S. Ports, along with assessment of zero- and near- zero emission fuel supplies at or near U.S. ports focused upon Hydrogen technologies. Initial market assessment for ocean going vessels for harbor support and ocean-going vessels is explored. An energy analysis is performed on the port system using a holistic approach and considering the port as an entire ecosystem that functions as a transportation and energy node. Presently, a comprehensive view is lacking for future analysis efforts, this analysis seeks to address this gap in data by evaluating four representative port types and the potential for utilizing hydrogen for the maritime industry. Every port is different, but broadly they could be bracketed into reference cases with scaling factors for the relative size of the port operations. These reference ports are for future use, potentially as baselines for analysis and development of demonstration programs. An equipment inventory for each reference port type (container, bulk, breakbulk, and inland waterway) is presented. A comparative analysis of fuel cell electric and battery electric equipment is conducted based on the following criteria: technology readiness level, refueling/charging time, operational range, energy consumption, and fuel cost savings compared to baseline internal combustion engine equipment. The tradeoffs and synergies between two alternative powertrains is highlighted. Based on energy and infrastructure analysis, average and high equipment utilization profiles across different port types is identified and quantified baseline fuel and electricity demand for various decarbonization scenarios. Based on the portfolio of equipment converted to fuel cell electric, the estimates of initial capital investment are provided for hydrogen refueling stations across ports. An energy demand model is developed that predicts well the all-electric cargo handling equipment annual energy consumption for ports with annual tonnage under 2 million twenty-foot equivalent units (TEUs). The model is a good rubric to follow for further energy demand models that can create a scalable solution to understand the energy needs of cargo handling equipment, whether they are all-electric, hydrogen fuel cell, or powered by another fuel-type. Zero and near-zero emission fuel supply at ports is evaluated looking into the characteristics of hydrogen, ammonia, and methanol as an alternative fuel, as well as the bunkering status. The readiness of reference ports to produce ammonia or methanol and bunker the fuel is examined based on the framework developed by the Global Maritime Forum and Rocky Mountain Institute.

08 HYDROGEN

Carbon Utilization and Storage Partnership of the Western United States

This technical report documents research conducted under DOE Award No. DE-FE0031837 focused on evaluating the feasibility of carbon capture, utilization, and storage (CCUS) systems in the central and western United States. The project integrated geologic characterization, reservoir simulation, infrastructure modeling, and economic analysis to assess CO₂ storage potential near industrial sources and develop strategies for transport and sequestration. The work included subsurface modeling, risk assessment, monitoring and verification (MRV) planning, and evaluation of regulatory pathways such as EPA Underground Injection Control (UIC) Class VI permitting and IRS 45Q tax credit eligibility. Results demonstrate the viability of multiple storage approaches, including saline formations, enhanced coalbed methane recovery, and basalt mineralization, supported by data-driven workflows and regional analyses. The project also produced permitting templates, technology transfer activities, and stakeholder engagement efforts to support deployment readiness. These findings contribute to the development of scalable, economically viable CCUS systems and provide a repeatable framework for future carbon management projects.

20 FOSSIL-FUELED POWER PLANTS

Observational Data for Next-Generation Climate Model Evaluation: Requirements, Considerations, and Best Practices

Climate model simulations are an important source of information about our planet’s climate system and also enable informed decision-making under different future scenarios. As a new archive of results from the next generation of climate models is anticipated to become available with the Coupled Model Intercomparison Project phase 7 (CMIP7), the need to develop efficient and robust methods to evaluate models is paramount. Observations are an integral part of model evaluation, providing a means to quantify and understand the degree to which climate models can faithfully reproduce Earth system processes. Such analysis is critical for constraining climate projections, identifying areas of focus for model development, and assisting analysts in deciphering the utility of models for specific applications. Observations of Earth system come from a diversity of sources, span different space–time domains, and are produced by different communities, and each dataset features different data structures and formats, metadata standards, and its own unique uncertainties. Uncertainties in an observational dataset may stem from gaps in temporal and spatial coverage, instrumentation errors, or assumptions in retrieval and processing methods. How then does one ensure that observational data are ready for use and utilized in the most appropriate way for robust, rapid, and routine climate model evaluation? The CMIP7 Model Benchmarking Task Team with input from the broader climate modeling, model evaluation, and observational data communities present a vision and considerations for best practices toward the optimal and appropriate use of observational data to support next-generation climate model evaluation.

Climate models

AI-Ready Control System for the Fermilab Accelerator Complex

Reliable, high-intensity operation of the Fermilab Accelerator Complex is critical to the success of the Long-Baseline Neutrino Facility and Deep Underground Neutrino Experiment. We describe the requirements and infrastructure necessary to support routine use of artificial intelligence and machine learning (AI/ML) in the accelerator control system. Three capabilities are identified: a machine learning operations (MLOps) framework standardizing the lifecycle of AI/ML automation from data management through deployment and monitoring; a data quality framework defining and enforcing standards required to build trustworthy AI/ML applications; and workflow integration with large language models to assist physicists, engineers, and operators with information retrieval, code development, and routine analysis. Use cases spanning beam diagnostics, beam control, and support system automation illustrate the technical requirements across the complex.

43 PARTICLE ACCELERATORS

Plant Bioengineering Atlas: A Knowledge Graph of Genes, DNA Constructs, and Plant Traits.

Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.

, Genes, DNA Constructs

Evaluating hydrogen permeation through medium-density and high-density polyethylene pipes to enable a hydrogen-compatible infrastructure

As the United States and the global community pursue energy abundance and resilience, hydrogen is becoming a critical energy carrier. The future of hydrogen transport depends on one thing: verifying that today’s pipelines are hydrogen ready. This study investigates the permeation behavior of medium-density polyethylene (MDPE) and high-density polyethylene (HDPE) pipelines under pure hydrogen gas environment at various pressures and temperatures. Using thermal desorption analysis (TDA) and direct pipe permeability measurements, hydrogen loss rates are quantified, revealing minimal permeation losses (< 10.5 g/km/day) under realistic operating conditions. Effects of temperature and pressure with Standard Dimension Ratios (SDR), were evaluated to determine their influence on permeation rates of MDPE and HDPE pipelines. In addition, ex-situ density and degree of crystallinity (DOC) of MDPE and HDPEs after hydrogen exposure were also investigated to understand pressure-dependency of polymer properties. The findings provide essential data for assessing polymer-based pipeline materials for safe and reliable hydrogen transport.

08 HYDROGEN

Automated Signal Timing Plan Reconstruction Using High-Resolution Event-Based Controller Data for Digital Twins

Transportation digital twins are essential tools for evaluating emerging technologies such as connected and automated vehicles, adaptive traffic signal control, and mobility optimization strategies. Realistic digital twins require accurate emulation of real-world signal controllers and detailed signal timing plans. However, signal timing plans are often unavailable or difficult to access, forcing researchers and modelers to rely on assumed fixed timings or halt their analysis. To overcome this challenge, we present a method that directly estimates signal timing plan parameters using high-resolution, event-based data from traffic signal controllers. The proposed method extracts key parameters, including cycle length, offset, phase sequence, coordinated phases, phase-specific minimum and maximum green durations, vehicle extensions, and splits under coordination. A rule-based deterministic signal timing reconstruction algorithm based on traffic signal operation rules, such as those outlined in the Signal Timing Manual, is developed and validated. We evaluate this method, which uses high-resolution controller event logs and verified signal timing plans, on 94 signalized intersections in Nashville, Tennessee, demonstrating their ability to generate accurate, simulation-ready signal timing plans for tools such as SUMO and Vissim.

Saroj, Abhilasha [ORNL] (ORCID:0000000191178063)