Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “standardized data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Guiding the choice of informatics software and tools for lipidomics research applications

Progress in mass spectrometry lipidomics has led to a rapid proliferation of studies across biology and biomedicine. These generate extremely large raw datasets requiring sophisticated solutions to support automated data processing. To address this, numerous software tools have been developed and tailored for specific tasks. However, for researchers, deciding which approach best suits their application relies on ad hoc testing, which is inefficient and time consuming. Here we first review the data processing pipeline, summarizing the scope of available tools. Next, to support researchers, LIPID MAPS provides an interactive online portal listing open-access tools with a graphical user interface. This guides users towards appropriate solutions within major areas in data processing, including (1) lipid-oriented databases, (2) mass spectrometry data repositories, (3) analysis of targeted lipidomics datasets, (4) lipid identification and (5) quantification from untargeted lipidomics datasets, (6) statistical analysis and visualization, and (7) data integration solutions. Detailed descriptions of functions and requirements are provided to guide customized data analysis workflows.

59 BASIC BIOLOGICAL SCIENCES↗

Small-angle X-ray and neutron scattering

Small-angle scattering (SAS) is a technique that is able to probe the structural organization of matter and quantify its response to changes in external conditions. X-ray and neutron scattering profiles measured from bulk materials or materials deposited at surfaces arise from nanostructural inhomogeneities of electron or nuclear density. Furthermore, the analysis of SAS data from coherent scattering events provides information about the length scale distributions of material components. Samples for SAS studies may be prepared in situ or under near-native conditions and the measurements performed at various temperatures, pressures, flows, shears or stresses, and in a time-resolved fashion. In this Primer, we provide an overview of SAS, summarizing the types of instrument used, approaches for data collection and calibration, available data analysis methods, structural information that can be obtained using the method, and data depositories, standards and formats. Recent applications of SAS in structural biology and the soft-matter and hard-matter sciences are also discussed.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Photovoltaic inverter-based quantification of snow conditions and power loss

Snow is a significant challenge for photovoltaic (PV) systems at northern latitudes, where the pace of deployment is rapid but snow-related power losses can exceed 30% of annual production. Accurate snow-related power loss estimation methods for utility-scale sites can support snow mitigation strategies, inform resource planning and validate predictive snow-loss models. This study builds on our previous work on inverter-based detection of snow, and its implications for utility-scale power production, by validating the accuracy of our snow-loss method across different PV sites and system designs and highlighting its value in bringing greater visibility to PV plant operations in winter. Our estimation method is both novel and scalable, requiring only standard monitoring data to correlate snow-related losses with meteorological data. As demonstrated here, our validation method involved three main steps: 1) estimation of performance losses for multiple systems by comparing measured inverter data to modeled data; 2) application of a detection framework to identify which performance losses are snow-related; and 3) comparison of snow-related losses among three utility-scale sites differing in tilt angle. Results show that utility-scale systems at higher tilt angles consistently shed snow more quickly/completely than their lower-tilt counterparts. Further, monthly and seasonal snow losses are inversely and non-linearly correlated with tilt angle when normalized for cumulative snowfall. These results are consistent with the findings of previous studies and support the broad applicability of this method to fixed-tilt utility-scale PV systems around the world that routinely experience snow-related performance losses.

Cooper, Emma C. (ORCID:0000000190554098)↗

Laser material interactions in tamped materials on picosecond time scales in aluminum

Here, a 100 ps laser is used to probe the pressure generation, depth of the non-solid ablator, and the non-linear optical effects through tamper materials. Samples consisted of an aluminum ablator with tampers of sapphire and coverslip glass. In general, the sapphire tamped sample achieves higher pressures at lower laser intensities as compared to the coverslip glass tamped sample. Attempts to model the details of this set of experimental data with standard available radiation coupled hydrodynamic codes make clear that more physics is needed in these simulations to accurately predict the impact of the tamper material on the pressure generation and the depth of non-solid aluminum.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Estimating internal moisture generation rates in recently constructed, occupied homes in the southeastern United States

Internal moisture generation (IMG), or moisture generated by building occupants via activities such as respiration, cooking, bathing, and cleaning, is a critical input required for design, analysis, and simulation of building enclosure and heating, ventilation, and air conditioning (HVAC) systems. Based on previously published values, ASHRAE Standard 160-2021 provides guidance for estimating IMG rates for moisture control design analysis, which is based on occupied home datasets collected in the 1980s and 1990s. Residential energy use simulation software also utilizes estimates for IMG for energy rating, energy analysis, and code compliance calculations, as specified in ANSI/RESNET/ICC Standard 301. Data quantifying IMG rates in newer homes is useful in determining the continued relevance of current design and simulation guidance, and whether the guidance represents conditions found in new housing stock. ASHRAE Research Project 1844, conducted by the Florida Solar Energy Center (FSEC), a research institute of the University of Central Florida, estimated IMG rates in newer occupied homes built since 2013 in the southeastern US using a moisture balance model approach. Furthermore, the project obtained occupied home data in cooperation with a US Department of Energy Building America research project to characterize indoor air quality (IAQ) in newer US homes, along with presence, functionality, and occupant use of control measures. Full-scale laboratory homes operating with known IMG rates were used to validate a moisture balance model and quantify the accuracy of estimates obtained when the model is applied to data from occupied homes with unknown IMG rates.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Scattering matrix pole expansions for complex wave numbers in R -matrix theory

In this followup article to Ducru et al., we establish new results on scattering matrix pole expansions for complex wave numbers in R-matrix theory. In the past, two branches of theoretical formalisms emerged to describe the scattering matrix in nuclear physics: R-matrix theory and pole expansions. The two have been quite isolated from one another. Recently, our study of Brune's alternative parametrization of R-matrix theory has shown the need to extend the scattering matrix (and the underlying R-matrix operators) to complex wave numbers. Two competing ways of doing so have emerged from a historical ambiguity in the definitions of the shift S and penetration P functions: the legacy Lane and Thomas's “force closure” approach versus analytic continuation (which is the standard in mathematical physics). The R-matrix community has not yet come to a consensus as to which to adopt for evaluations in standard nuclear data libraries, such as ENDF. Here, in this article, we argue in favor of analytic continuation of R-matrix operators. We bridge R-matrix theory with the Humblet-Rosenfeld pole expansions, and discover new properties of the Siegert-Humblet radioactive poles and widths, including their invariance properties to changes in channel radii a c . We then show that analytic continuation of R-matrix operators preserves important physical and mathematical properties of the scattering matrix—canceling spurious poles and guaranteeing generalized unitarity—while still being able to close channels below thresholds.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Characterizing Sub-Cohorts via Data Normalization and Representation Learning

The process of identifying a cohort of interest is a very challenging task. It requires manually inspecting many patient records of complex structure that might include medical coding errors and missing data. This paper presents a computational pipeline for refining the process of cohort selection based on medical concepts recorded in the electronic health records (EHRs). The pipeline extracts EHR data for a given cohort and normalizes this data using standard vocabularies. Then a stacked denoising autoencoder is used to embed the normalized patient vectors in a low dimensional space, where the patients are subsequently clustered into sub-cohorts. The goal is to represent the cohort in a standard format and abstract variants of sub-populations. As a use-case, we applied the pipeline to 1.8 million Veterans diagnosed with major depressive disorder (MDD), and identified four meaningful sub-cohorts using the features learned by the autoencoder. Then, each sub-cohort was explored using a set of keywords for interpretation.

Rush III, Everett↗

Attention-Augmented Parametric Kernel Graph Neural Network (APKGNN) for Node Classification

We present a new graph neural network, the Attention-based Parametric-Kernel augmented Graph Neural Network (APKGNN), developed for node classification tasks. Despite extensive work on modeling multi-faceted relationships between connected nodes of a graph, the effect of attention on edge features mapped to relationships has not yet been analyzed through learning representation. This study derives such an attention vector by first calculating node features corresponding to endpoints of an edge and then aggregating these with extracted local intrinsic patches of a given graph to generate augmented local patch vectors. This process uses a parametric kernel based on Gaussian mixture models (GMMs) to embed local neighborhoods of the graph in local patches. The patch vectors then convolve with the above node features to produce an updated node representation. We show that this new learning representation (APKGNN) achieves higher node classification accuracy on tasks - both standard benchmarks (Cora, PubMed, Citeseer) and new experimental short text corpora where nodes correspond to text documents and words. This implementation of the GNN convolution layer outperforms state-of-the-art (SOTA) algorithms, achieving higher training, validation, and test accuracy by a significant margin on three standard benchmark data sets under both SOTA experimental settings and those for new testbeds.

Bose, Avishek↗

Model-based interface design for smart field-device integration

Operational complexity is ever-increasing for electric utilities that face challenges including integration of DERs, customer expectation of energy choices, the proliferation of non-utility-owned resources, new business models with energy service providers, and new technology with IT/OT convergence. To maintain and improve the quality of operations, planning, and decision-making in general, utilities need to manage and navigate the complexity. Managing complexity requires a modular, scalable, and flexible solution. Connecting large amounts of DERs and introducing new services requires increased grid control and evolving applications. In this paper, we show an approach utilizing model-based standardized interfaces that simplifies integration and deployment of new algorithms and smart field devices for interoperability across legacy or new systems. A modular design is presented for a reference implementation of widely used DNP3 and IEEE 2030.5 interfaces within an open-source, standards-based data integration platform for integration of smart field devices with independently developed, best-of-breed applications.

Model-driven development, system integration, smar↗

Machine Learning Based Resilience Testing of an Address Randomization Cyber Defense

Moving target defenses (MTDs) are widely used as an active defense strategy for thwarting cyberattacks on cyber-physical systems by increasing diversity of software and network paths. Recently, machine Learning (ML) and deep Learning (DL) models have been demonstrated to defeat some of the cyber defenses by learning attack detection patterns and defense strategies. It raises concerns about the susceptibility of MTD to ML and DL methods. Here, in this article, we analyze the effectiveness of ML and DL models when it comes to deciphering MTD methods and ultimately evade MTD-based protections in real-time systems. Specifically, we consider a MTD algorithm that periodically randomizes address assignments within the MIL-STD-1553 protocol—a military standard serial data bus. Two ML and DL-based tasks are performed on MIL-STD-1553 protocol to measure the effectiveness of the learning models in deciphering the MTD algorithm: 1) determining whether there is an address assignments change i.e., whether the given system employs a MTD protocol and if it does 2) predicting the future address assignments. The supervised learning models (random forest and k-nearest neighbors) effectively detected the address assignment changes and classified whether the given system is equipped with a specified MTD protocol. On the other hand, the unsupervised learning model (K-means) was significantly less effective. The DL model (long short-term memory) was able to predict the future addresses with varied effectiveness based on MTD algorithm's settings.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Observation of four-top-quark production in the multilepton final state with the ATLAS detector

This paper presents the observation of four-top-quark ($t$$\overline{t}$$t$$\overline{t}$) production in proton-proton collisions at the LHC. The analysis is performed using an integrated luminosity of 140 fb -1 at a centre-of-mass energy of 13 TeV collected using the ATLAS detector. Events containing two leptons with the same electric charge or at least three leptons (electrons or muons) are selected. Event kinematics are used to separate signal from background through a multivariate discriminant, and dedicated control regions are used to constrain the dominant backgrounds. The observed (expected) significance of the measured $t$$\overline{t}$$t$$\overline{t}$ signal with respect to the standard model (SM) background-only hypothesis is 6.1 (4.3) standard deviations. The $t$$\overline{t}$$t$$\overline{t}$ production cross section is measured to be ${22.5}^{+6.6}_{-5.5}$, consistent with the SM prediction of 12.0 ± 2.4 fb within 1.8 standard deviations. Data are also used to set limits on the three-top-quark production cross section, being an irreducible background not measured previously, and to constrain the top-Higgs Yukawa coupling and effective field theory operator coefficients that affect $t$$\overline{t}$$t$$\overline{t}$ production.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

HPXML Version Translator

HPXML is a consensus data transfer standard for residential buildings. Over several years, multiple versions of the standard have been released. This tool accepts an HPXML file in an older version and translates it to a newer version.

Merket, Noel↗

Python wrapper library and analysis functions for Geotab Altitude API [SWR-24-77]

This software library serves as a Python wrapper for Geotab's Altitude API. It streamlines querying of the API, converts loosely structured API outputs into a standardized tabular data format, and enables analysis of the resulting data tables. It also includes example notebooks showing how to use the library.

Bruchon, Matthew↗

lanl-ansi/MG-RAVENS

The MG-RAVENS project with the DOE Office of Electricity Microgrid R&D Program is a project to develop a completely free, open-source data exchange standard (API) for the Department of Energy, targeted at software tools related to infrastructure modeling, particularly the modeling of microgrids and electric power distribution systems that are created with funding from the Microgrid R&D Program. This software produces formal definitions of an API, documentation, contains supporting functions for parsing, validating, etc., and will contain examples of workflows enabled by the developed API.

Fobes, David M↗

GeoCricket

SAND2025-12229O Geospatial Critical Infrastructure and Census Data Stockpile Tool (GeoCricket) is a set of functions that collect critical infrastructure and census data for use in the Resilient Node Cluster Analysis Tool (ReNCAT) and Quantum Geographic Information System Social Burden Calculator. It can also act to inform other place-based work. The code queries public-facing Representational State Transfer (REST) servers to collect geospatial data related to a specific area. It then exports that data as standard geographic information system file types or as a .csv file. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Haines, John↗

Single-shot electro-optic sampling with arbitrary terahertz polarization

With the recent development of diversity electro-optic sampling (DEOS), significant progress has been made in the range of applicability of single-shot EOS measurements, allowing broadband THz waveforms to be captured in a single shot over large temporal windows. In addition to the decrease in acquisition time compared to standard multishot data acquisition, this technique allows measurements on systems far from equilibrium with large shot-to-shot noise or with irreversible or poorly repeatable dynamics. Although DEOS has been demonstrated and verified for linearly polarized THz waveforms, we investigate the effects resulting from the presence of a secondary polarization component. This imposes new challenges for accurate waveform reconstruction, and opens the opportunity to measure out complex polarization states such as arbitrary elliptically polarized THz field. We demonstrate a single-shot diversity-electro-optic-sampling-based approach to capture both x- and y-THz fields simultaneously with a single (110)-cut EO crystal for THz polarimetry and ellipsometry over a wide range of frequencies.

Lenz, Maximilian [University of California, Los An↗

ESS-DIVE Reporting Format for Amplicon Abundance Table

While standardized sequencing data is available in public repositories and efforts such as MIxS for common sample collection and processing metadata are well established, the lack of common bioinformatic processing metadata has hindered the ability to do large-scale metaanalyses and the potential for data re-use by non-experts such as ecosystem, watershed, or earth system modelers. To address this need for Department of Energy researchers, we have developed an amplicon reporting format which captures both sample preparation and bioinformatic processing metadata and stores processed amplicon data as a paired abundance table and sequencing file to maximize the potential for re-use of these data. To aid in the adoption of accessible and reproducible analysis workflows, this reporting format was developed in concert with amplicon functionality within the Department of Energy’s Systems Biology Knowledgebase (KBase) to ensure common data and metadata requirements and facilitate seamless transfer between these platforms.This dataset contains support documentation for the amplicon reporting format (README.md and instructions.md), templates for both bioinformatic and sequencing metadata (amplicon_bioinformatic_metadata_template_2021_10_03.csv and amplicon_sequencing_metadata_template_2021_10_03.csv), a crosswalk indicating how this reporting format relates to the current MIxS format (ESSDIVE-MIxS_crosswalk.csv), a list of available instrument terms (amplicon_seq_instrument_terms_2021_10_03.csv), a map between QIIME2 parameter settings and metadata fields (amplicon_qiime2_plugin_metadata_map.csv), a data dictionary (amplicon_CSV_dd.csv), and file-level metadata (amplicon_FLMD.csv).

54 ENVIRONMENTAL SCIENCES↗

Observed and Imputed Volumetric Soil Water Content Timeseries for the New Mexico Elevation Gradient

Reliable soil water content (SWC) data are essential for understanding dryland ecosystem dynamics, but high-frequency SWC sensors often fail, creating gaps in critical datasets. To address this, we developed a Bayesian mixture model that imputes missing SWC using both linear interpolation and an ecosystem water balance model (SOILWAT2), tested across six AmeriFlux eddy covariance tower sites in the New Mexico Elevation Gradient, demonstrating its effectiveness in reconstructing SWC patterns while providing insights into the factors driving SWC variability. Daily volumetric soil water content (SWC) data are provided as csv-formatted spreadsheets for the six AmeriFlux sites (US-Seg, US-Ses, US-Wjs, US-Mpi, US-Vcp, and US-Vcs). For each site there is an observed SWC file (site_SWC_gapfill.csv) and a file that contains imputed SWC (imputed_SWC_site.csv). The observed SWC files contain temperature corrected sensor values, tower precipitation data, as well as outputs from SOILWAT2 simulations that were used to impute SWC. The imputed files contain the original observed SWC values and the imputed missing SWC values. When SWC was missing from the original data, the missing value was imputed based on the Bayesian imputation mixture model. The posterior mean of all imputed values is reported as "mean_X". When the observed SWC was NOT missing, mean_X = observed SWC value (original data). The standard deviation, 2.5th percentile and the 97.5th percentile for the imputed values are also reported in the imputed files. There are readme text files for each file type explaining the contents of each column.

54 ENVIRONMENTAL SCIENCES↗