Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data systems standards”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Generalizable Web User Interface for Scalable and Streamlined Deployment of Building Energy Management Systems in Small and Medium-Sized Commercial Buildings

Small and medium-sized commercial buildings (SMCBs) comprise 94% of US commercial buildings yet face significant barriers to implementing building energy management systems despite advances in smart device technology. Existing solutions present critical limitations: cloud-based API solutions simplify deployment but create vendor lock-in constraints; commercial integrated software solutions ensure compatibility via standardized protocols but require substantial cost and technical expertise; open-source IoT platforms offer cost-effective vendor independence but provide insufficient standardized protocol support for commercial building automation. This research presents a generalizable web user interface framework that bridges the gap between evolving smart device capabilities and lagging software infrastructure for SMCBs. The proposed system integrates VOLTTRON open-source middleware with an automated configuration converter that transforms unified specifications written in YAML, a human-readable data-serialization format, into system-specific files, streamlining manual setup processes. The vendor-agnostic architecture supports industry-standard protocols (BACnet and Modbus) and semantic building models while providing adaptive web interfaces that dynamically adjust to various building configurations. Demonstrations through simulation-based testing and a field deployment show automatic interface adaptation across heterogeneous HVAC systems and multizone monitoring. The automated configuration converter also substantially reduces labor-intensive setup.

Chung, Jihoon [ORNL] (ORCID:0000000184880815)↗

The Phase-2 Upgrade of the CMS Data Acquisition

The High Luminosity LHC (HL-LHC) will start operating in 2027 after the third Long Shutdown (LS3), and is designed to provide an ultimate instantaneous luminosity of 7:5 × 10$^{34}$ cm$^{-2}$ s$^{-1}$, at the price of extreme pileup of up to 200 interactions per crossing. The number of overlapping interactions in HL-LHC collisions, their density, and the resulting intense radiation environment, warrant an almost complete upgrade of the CMS detector. The upgraded CMS detector will be read out by approximately fifty thousand highspeed front-end optical links at an unprecedented data rate of up to 80 Tb/s, for an average expected total event size of approximately 8 - 10 MB. Following the present established design, the CMS trigger and data acquisition system will continue to feature two trigger levels, with only one synchronous hardware-based Level-1 Trigger (L1), consisting of custom electronic boards and operating on dedicated data streams, and a second level, the High Level Trigger (HLT), using software algorithms running asynchronously on standard processors and making use of the full detector data to select events for offline storage and analysis. The upgraded CMS data acquisition system will collect data fragments for Level-1 accepted events from the detector back-end modules at a rate up to 750 kHz, aggregate fragments corresponding to individual Level- 1 accepts into events, and distribute them to the HLT processors where they will be filtered further. Events accepted by the HLT will be stored permanently at a rate of up to 7.5 kHz. This paper describes the baseline design of the DAQ and HLT systems for the Phase-2 of CMS.

Badaro, Gilbert↗

Cataloging Legacy Data from the Tritium Systems Test Assembly Program

The Tritium Systems Test Assembly (TSTA) at Los Alamos National Laboratory, operational from 1984 to 2001, was critical in advancing fusion fuel cycle technologies, including tritium storage, gas separation, and pumping. TSTA’s contributions, particularly in safe tritium operations, have influenced subsequent fusion projects. This paper discusses the ongoing effort to digitize and catalog TSTA’s historical data to create a searchable resource for the fusion research community. While the long-term objective is to develop a relational database for structured data management, the project remains in the early phase, with current efforts focused on scanning and indexing physical documents. Initial plans for database implementations are also presented, outlining key considerations for structure, query indexing, and standardization. As digitization progresses, future discussions will refine these implantation details to ensure an efficient and comprehensive system. This initiative aims to preserve critical legacy data, enhance the design of tritium system facilities, and support the next generation of fusion energy research.

42 ENGINEERING↗

Remote Instrumentation and Data Acquisition

This poster outlines the development and implementation of a remote data acquisition system for waveform analysis using a Rohde & Schwarz oscilloscope. The project involved capturing waveform data, and transferring it to a local machine for visualization and analysis. The core logic was developed in C++ with a focus on object oriented programming and the use of polymorphism so the main application can interact with any instrument without knowing its exact type, simplifying the overall logic and making it easier to add or swap out components without changing the rest of the codebase.. The system issues Standard Commands for Programmable Instruments (SCPI) via a socket connection and parses the oscilloscope s ASCII waveform data. The C++ application was containerized using Docker for ease of portability, and reproducibility. Emphasis was placed on secure networking practices, error handling, and effective data capture. The report describes the technical steps taken, challenges encountered, and future work, providing insight into the practical integration of hardware interfacing with remote computational environments.

Parikh, Jaymil [Illinois U., Urbana]↗

Remote Instrumentation and Data Acquisition: An Internship Research Report

This report outlines the development and implementation of a remote data acquisition system for waveform analysis using a Rohde & Schwarz oscilloscope. The project involved capturing waveform data, and transferring it to a local machine for visualization and analysis. The core logic was developed in C++ with a focus on object oriented programming and the use of polymorphism so the main application can interact with any instrument without knowing its exact type, simplifying the overall logic and making it easier to add or swap out components without changing the rest of the codebase.. The system issues Standard Commands for Programmable Instruments (SCPI) via a socket connection and parses the oscilloscope’s ASCII waveform data. The C++ application was containerized using Docker for ease of portability, and reproducibility. Emphasis was placed on secure networking practices, error handling, and effective data capture. The report describes the technical steps taken, challenges encountered, and lessons learned, providing insight into the practical integration of hardware interfacing with remote computational environments.

Parikh, Jaymil [Fermilab]↗

Infrared Cloud Imager Instrument Intercomparison Report

The Infrared Cloud Imager Instrument Intercomparison was a guest instrument deployment by NWB Sensors to the U.S. Department of Energy’s Atmospheric Radiation Measurement (ARM) User Facility observatory on the Southern Great Plains (SGP) between May 18 and December 12, 2023. NWB Sensors is a company that has developed a commercially available infrared cloud imager (ICI). The ICI provides radiometrically calibrated, full-sky images of the downwelling infrared radiance in the 7.3-14 µm band. In addition, it provides cloud radiance as the residual between the observed radiance and the modeled cloud-free radiance as well as derived cloud products. The instrument is used in applications that require consistent detection of clouds across day and night. For more information, consult the instrument's webpage. The primary goal of the deployment was to validate the radiometric accuracy of the ICI. The ICI uses a proprietary calibration method to convert the raw data from its infrared camera into downwelling radiance. Unlike similar instruments, the system does not have an onboard blackbody calibration standard. Instead, NWB Sensors characterizes each ICI camera individually in an environmental chamber while looking at a blackbody standard. The resulting (proprietary) calibration is used operationally in the instrument and has been demonstrated to be stable over long periods. To validate the radiometric products from the ICI, an intercomparison between the ICI data products and those from ARM’s atmospheric emitted radiance interferometer (AERI) was made. The AERI is a best-in-class instrument for measuring downwelling infrared radiance (Gero et al. 2025). A weighted integration of the AERI’s spectral radiances across the ICI’s camera response was performed. The resulting radiance (herein called the AERI radiance) was directly compared to the zenith radiance concurrently observed by the ICI. The results of these comparisons are reported in the next section of this report.

54 ENVIRONMENTAL SCIENCES↗

Evaluating the Incident Energy of Arcs in Photovoltaic DC Systems: Comparison Between Calculated and Experimental Data

Solar Photovoltaic (PV) systems have permeated the energy generation world at a very high rate, some of the safety codes and standards are still lagging in accurately assessing the hazards and risks associated with PV array arcing energies. Safety professionals and maintenance workers using NFPA 70E have utilized the Doan, Stokes & Oppenlander or Enrique models, meant to determine arc energies in DC power systems using the maximum power method. These methods may lead to an overestimate of energy available in PV systems during a fault. Since PV modules/arrays are non-linear, current limited DC devices, some of these calculation methods may not accurately predict fault energy. This paper will validate current arc energy models for PV systems by comparing experimental and calculated data. Additionally, this data will help modify the current NFPA 70E models related to smaller solar arrays. Understanding where the real safety threshold for DC arc flash in PV systems exists will help maintenance and safety professionals better prepare for a variety of work related activities. This paper will analyze real arc data taken for PV systems <1000VDC and <60amps and compare this to the calculated incident energy models, to include 70E. Thus, using these comparisons, it may be possible to reduce the safety hazard severity and thus relax the PPE requirements for installation and maintenance crews.

14 SOLAR ENERGY↗

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence↗

Developing an Interactive Landscape for Mobility Resources: Preprint

As the world continues to be increasingly driven by data, the ways researchers and professionals sort and collect this data is critical. In the world of mobility data, new levels of data from public transportation systems, location services, and other means are being lost due to how little organization exists. Much of the data is proprietary, and there are few if any de jure or even de facto standards connecting data. There is also little knowledge about the gaps that exist in the data. In this project, we created an interactive landscape where mobility resources are categorized and organized in an easy to use, living document. We made this landscape with open-source code from the CNCF Cloud Native Landscape and repurposed it to the mobility data's needs. Additionally, unlike previous sources that organize mobility data, this document can be updated through GitHub by those in the field to keep its sources relevant. Following the creation of a beta version of the landscape, we conducted several interviews with industry researchers and professionals to ensure the landscape would be useful. The result is an online hub where mobility researchers and resource creators can easily access research and collaborate.

ADVANCED PROPULSION SYSTEMS↗

Machine Learning Based Resilience Testing of an Address Randomization Cyber Defense

Moving target defenses (MTDs) are widely used as an active defense strategy for thwarting cyberattacks on cyber-physical systems by increasing diversity of software and network paths. Recently, machine Learning (ML) and deep Learning (DL) models have been demonstrated to defeat some of the cyber defenses by learning attack detection patterns and defense strategies. It raises concerns about the susceptibility of MTD to ML and DL methods. Here, in this article, we analyze the effectiveness of ML and DL models when it comes to deciphering MTD methods and ultimately evade MTD-based protections in real-time systems. Specifically, we consider a MTD algorithm that periodically randomizes address assignments within the MIL-STD-1553 protocol—a military standard serial data bus. Two ML and DL-based tasks are performed on MIL-STD-1553 protocol to measure the effectiveness of the learning models in deciphering the MTD algorithm: 1) determining whether there is an address assignments change i.e., whether the given system employs a MTD protocol and if it does 2) predicting the future address assignments. The supervised learning models (random forest and k-nearest neighbors) effectively detected the address assignment changes and classified whether the given system is equipped with a specified MTD protocol. On the other hand, the unsupervised learning model (K-means) was significantly less effective. The DL model (long short-term memory) was able to predict the future addresses with varied effectiveness based on MTD algorithm's settings.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Bioenergy Underground: Challenges and opportunities for phenotyping roots and the microbiome for sustainable bioenergy crop production

Abstract Bioenergy production often focuses on the aboveground feedstock production for conversion to fuel and other materials. However, the belowground component is crucial for soil carbon sequestration, greenhouse gas fluxes, and ecosystem function. Roots maximize feedstock production on marginal lands by acquiring soil resources and mediating soil ecosystem processes through interactions with the microbial community. This belowground world is challenging to observe and quantify; however, there are unprecedented opportunities using current methodologies to bring roots, microbes, and soil into focus. These opportunities allow not only breeding for increased feedstock production but breeding for increased soil health and carbon sequestration as well. A recent workshop hosted by the USDOE Bioenergy Research Centers highlighted these challenges and opportunities while creating a roadmap for increased collaboration and data interoperability through standardization of methodologies and data using F.A.I.R. principles. This article provides a background on the need for belowground research in bioenergy cropping systems, a primer on root system properties of major U.S. bioenergy crops, and an overview of the roles of root chemistry, exudation, and microbial interactions on sustainability. Crucially, we outline how to adopt standardized measures and databases to meet the most pressing methodological needs to accelerate root, soil, and microbial research to meet the pressing societal challenges of the century.

09 BIOMASS FUELS↗

ARM Aerial Facility (AAF) - Unmanned Aircraft Systems, Cloud Droplet Probe with QC and lat/lon/alt

The Cloud Droplet Probe (CDP) is designed to measure cloud droplet size distribution from 2 µm to 50 µm. The CDP and an appropriate data system can also calculate various other parameters including particle concentrations, effective diameter (ED), Median Volume Diameter (MVD), and Liquid Water Content (LWC). The b1 level adds standard quality control flags, and merges in flight navigation data.

54 ENVIRONMENTAL SCIENCES↗

Importance of Standardizing Analytical Characterization Methodology for Improved Reliability of the Nanomedicine Literature

Understanding the interaction between biological structures and nanoscale technologies, dubbed the nano-bio interface, is required for successful development of safe and efficient nanomedicine products. The lack of a universal reporting system and decentralized methodologies for nanomaterial characterization have resulted in a low degree of reliability and reproducibility in the nanomedicine literature. As such, there is a strong need to establish a characterization system to support the reproducibility of nanoscience data particularly for studies seeking clinical translation. Here, we discuss the existing key standards for addressing robust characterization of nanomaterials based on their intended use in medical devices or as pharmaceuticals. We also discuss the challenges surrounding implementation of such standard protocols and their implication for translation of nanotechnology into clinical practice. We, however, emphasize that practical implementation of standard protocols in experimental laboratories requires long-term planning through integration of stakeholders including institutions and funding agencies.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

DOE BSSD Performance Management Metrics Report Q1

Microbes play key roles in our biosphere, from driving global nutrient cycling to impacting plant, animal and human health and disease. Complex data from microbial genomes, proteins, and metabolites provide a window into these tiny engines that drive life on our planet. Yet these data are dispersed among researchers’ laboratories and various repositories, making it difficult to access. This calls for new ways of managing data, improving data interoperability, advancing community standards, and creating an infrastructure where data are shared efficiently. We have built the National Microbiome Data Collaborative (NMDC) to advance how scientists create, use, and reuse data to redefine the way we understand and harness the power of microbes. The vision of the National Microbiome Data Collaborative (NMDC) is to drive a microbiome data sharing network connecting data, people, and ideas to advance microbiome innovation and discovery. The NMDC was launched in 2019 and brought together DOE National Laboratories to collaborate across resources, capabilities, and expertise. The NMDC team was strategically assembled to include software developers, microbial researchers, metadata experts, and multi-omics specialists. The diversity of the NMDC team reflects the inherently interdisciplinary nature of microbiome science, and we leverage the strengths of the DOE National Laboratory system. Towards BER’s goal of advancing an iterative systems biology approach to the understanding of microbial genomes, the NMDC serves as a foundation for infrastructure, data standards, and community building. Together with the flagship DOE User Facilities, the Joint Genome Institute (JGI) and the Environmental Molecular Sciences Laboratory (EMSL), we are developing core capabilities in metadata standards for environmental descriptors and sample handling and processing; standardized bioinformatic workflows; an interface for data search and access; and robust community engagement activities. The NMDC production platform supports long-term data infrastructure and community building for BER’s bioenergy and environmental research goals. Our approach leverages lessons learned and an ambitious framework for collaborative, interdisciplinary data infrastructure to support microbiome research. The NMDC supports data, information, and knowledge access through three defined software tools – the Submission Portal, NMDC EDGE, and the Data Portal – driven by community needs. Herein, we describe the value proposition for the microbiome research community, our overarching strategy, and challenges and opportunities for developing the NMDC as both an infrastructure and community engagement program.

59 BASIC BIOLOGICAL SCIENCES↗

Bridging the Gap on Data and Analysis for Distribution System Planning: Information That Utilities Can Provide Regulators, State Energy Offices and Other Stakeholders

Electric utilities conduct planning annually to ensure their distribution system meets technical standards, policies, and regulations; addresses forecasted grid conditions; satisfies customer needs; and advances utility priorities. The plan identifies grid deficiencies, analyzes potential solutions, and prioritizes capital investments and other expenditures. About 20 U.S. states and jurisdictions require regulated utilities to file some type of distribution system plan with the public utility commission for review. Requirements for sharing distribution system data and analyses vary widely, from few specific requirements to a detailed list of information that must be provided. While utilities conduct extensive analysis to develop distribution system plans, in most jurisdictions regulators and stakeholders do not know what data are available and how the utility uses the data in planning and investing. This report aims to bridge the gap by increasing understanding of the types of data and analyses utilities employ to develop distribution system plans and how the information affects their decision-making. The report describes information that states and stakeholders can ask for related to 11 data categories: -Forecasting loads and distributed energy resources (DERs) -Scenario analysis -Worst-performing circuits -Asset management strategy -Hosting capacity analysis -Value of DERs -Grid needs assessment -Cost-effectiveness framework for investments -Distribution system investment strategy and implementation -Geotargeted programs -Non-wires alternatives procurements.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Physics with high-luminosity proton-nucleus collisions at the LHC

The physics case for the operation of high-luminosity proton-nucleus (pA) collisions at the CERN LHC is reviewed. The collection of $\mathcal{O}$(1–10 pb −1 ) of proton-lead (pPb) collisions at the LHC will provide unique physics opportunities in a broad range of topics including proton and nuclear parton distribution functions (PDFs and nPDFs), generalised parton distributions (GPDs), transverse momentum dependent PDFs (TMDs), low-x quantum chromodynamics and parton saturation, hadron spectroscopy, baseline studies for quark-gluon plasma and parton collectivity, double and triple parton scatterings, photon–photon collisions, and physics beyond the Standard Model; which are not otherwise as clearly accessible by exploiting data from any other colliding system at the LHC. This report summarises the accelerator aspects of high-luminosity pA operation at the LHC, as well as each of the physics topics outlined above, including the relevant experimental measurements that motivate much larger pA datasets than collected to date.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Ectomycorrhizal effects on decomposition are highly dependent on fungal traits, climate, and litter properties: A model-based assessment. Dataset.

To simulate the effects of mycorrhizal fungi on soil organic matter cycling, we incorporated mycorrhizal processes into the Carbon, Organisms, Rhizosphere, and Protection in the Soil Environment (CORPSE) model to develop a new soil model Myco-CORPSE. The new model was calibrated and evaluated against soil measurements taken at temperate forests in New Hampshire (NH) and Georgia (GA). A series of scenario analysis were also conducted to explore the conditions under which ectomycorrhizal (ECM) N acquisition processes can induce different soil C accumulation in ECM systems compared to arbuscular (AM) systems.In this data package, we included:-The Python codes of the standard Myco-CORPSE model we developed: "Standard Myco_CORPSE python codes.zip". The main program is the "gradient_sim.py" which calculates the bulk soil microbes and CN content along a user defined gradient of clay, soil temperature, soil moisture and mycorrhizal dominance, and relies on two subprograms "CORPSE_deriv.py" and "CORPSE_integrate.py". "CORPSE_deriv.py" calculated the changes in all simulated soil stock within every time step and "CORPSE_integrate.py" integrate the changes in all simulated soil stock within simulated time period. The program "Plot.py" is used to plot the major outputs produced by the main program "gradient_sim.py".-The modified Python codes of Myco-CORPSE models with site-level environmental inputs (in NH and GA) used to conduct simulations in NH and GA sites: "NH_GA model simulations.zip". -The Python codes used to evaluate the Myco-CORPSE simulation outputs in NH and GA sites against site-level measurements: "Plot NH_GA simulation against measurements.zip". It includes both the evaluation Python code, the model outputs on NH and GA sites, and the measured soil properties in both sites.-The modified Python codes of Myco-CORPSE models "Scenario analysis_model simulations.zip" that is used to conduct scenario analysis of how different litter properties, mycorrhizal fungal traits, climate, and seasonal variation in temperature and vegetation phenology impact the mycorrhizal effects on soil CN properties. The sub file folder "Scenario analysis_litter traits" contains the codes for scenario analysis of different litter properties; The sub file folder "Scenario analysis_ECM types" contains the codes for scenario analysis of different ECM fungal traits; The sub file folder "Scenario analysis_climate&seasonality" contains the codes for scenario analysis of different climate and seasonalities;-"Scenario analysis_model results and plotting codes.zip" contains all the output files from the the scenario analysis of Myco-CORPSE model as described above and the plotting codes used to the generate the heatmaps and scatterplots shown in the manuscript "Ectomycorrhizal effects on decomposition are highly dependent on fungal traits, climate, and litter properties: A model-based assessment"The majority of the model outputs did not have specific geographic information or temporal coverage because the analysis we conducted are mainly hypothetical model simulations. We only provided geographic description, coordinates and temporal coverage for those soil measurements which we used for model evaluations (included in the "Plot NH_GA simulation against measurements.zip").

54 ENVIRONMENTAL SCIENCES↗

ATLAS-MAP: An Automated Test Station for Gated Electronic Transport Measurements

The diversification of electronic materials in devices provides a strong incentive for methods to rapidly correlate device performance with fabrication decisions. In this work, we present a low-cost automated test station for gated electronic transport measurements of field-effect transistors. Utilizing open-source PyMeasure libraries for transparent instrument control, the “ATLAS-MAP” system serves as a customizable interface between sourcemeters and samples under test and is programmed to conduct transfer curve and van der Pauw methods with static and sweeping gate voltages. Zinc oxide transistors of variable thickness (5, 10, and 20 nm) and channel size (50 μm to 3 mm, of equal length and width) were fabricated to validate the design. Standardization of testing procedures and raw data formatting enabled automated data analysis. A detailed list of parts and code files for the system are provided.

36 MATERIALS SCIENCE↗