Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data enrichment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Enriching Load Data Using Micro-PMUs and Smart Meters

In modern distribution systems, load uncertainty can be fully captured by micro-PMUs, which can record high-resolution data; however, in practice, micro-PMUs are installed at limited locations in distribution networks due to budgetary constraints. In contrast, smart meters are widely deployed but can only measure relatively low-resolution energy consumption, which cannot sufficiently reflect the actual instantaneous load volatility within each sampling interval. In this paper, we have proposed a novel approach for enriching load data for service transformers that only have low-resolution smart meters. The key to our approach is to statistically recover the high-resolution load data, which is masked by the low-resolution data, using trained probabilistic models of service transformers that have both high- and low-resolution data sources, i.e., micro-PMUs and smart meters. The overall framework consists of two steps: first, for the transformers with micro-PMUs, a Gaussian Process is leveraged to capture the relationship between the maximum/minimum load and average load within each low-resolution sampling interval of smart meters; a Markov chain model is employed to characterize the transition probability of known high-resolution load. Next, the trained models are used as teachers for the transformers with only smart meters to decompose known low-resolution load data into targeted high-resolution load data. The enriched data can recover instantaneous load uncertainty and significantly enhance distribution system observability and situational awareness. Here, we have verified the proposed approach using real high- and low-resolution load data.

24 POWER TRANSMISSION AND DISTRIBUTION↗

NP-MRD: the Natural Products Magnetic Resonance Database

The Natural Products Magnetic Resonance Database (NP-MRD) is a comprehensive, freely available electronic resource for the deposition, distribution, searching and retrieval of nuclear magnetic resonance (NMR) data on natural products, metabolites and other biologically derived chemicals. NMR spectroscopy has long been viewed as the ‘gold standard’ for the structure determination of novel natural products and novel metabolites. NMR is also widely used in natural product dereplication and the characterization of biofluid mixtures (metabolomics). All of these NMR applications require large collections of high quality, well-annotated, referential NMR spectra of pure compounds. Unfortunately, referential NMR spectral collections for natural products are quite limited. It is because of the critical need for dedicated, open access natural product NMR resources that the NP-MRD was funded by the National Institute of Health (NIH). Since its launch in 2020, the NP-MRD has grown quickly to become the world's largest repository for NMR data on natural products and other biological substances. It currently contains both structural and NMR data for nearly 41,000 natural product compounds from >7400 different living species. All structural, spectroscopic and descriptive data in the NP-MRD is interactively viewable, searchable and fully downloadable in multiple formats. Extensive hyperlinks to other databases of relevance are also provided. The NP-MRD also supports community deposition of NMR assignments and NMR spectra (1D and 2D) of natural products and related meta-data. The deposition system performs extensive data enrichment, automated data format conversion and spectral/assignment evaluation.

59 BASIC BIOLOGICAL SCIENCES↗

Super enrichments of Fe-group nuclei in solar flares and their association with large He-3 enrichments

Data on solar flares and perodic particle intensity enhancements in the energy range from 1 to 20 MeV/n are examined. It is found that: (1) Fe/He-4 ratios range from about 1 to 1000 times the solar ratio of 0.0004; (2) these high ratios mitigate against extended storage and large amounts of nuclear processing; (3) the CNO/He-4 ratio has a much smaller range of variability and a mean value of 0.02; (4) large He-3 and Fe enrichments are strongly associated, but not on a one-to-one basis; (5) large Fe enhancements sometimes occur without correspondingly large He-3 enrichments; and (6) none of the models so far advanced adequately explains the observed He-3 and heavy-nucleus enrichments.

Anglin, J. D.↗

Benefits and Limits of Phasing Alleles for Network Inference of Allopolyploid Complexes

Abstract Accurately reconstructing the reticulate histories of polyploids remains a central challenge for understanding plant evolution. Although phylogenetic networks can provide insights into relationships among polyploid lineages, inferring networks may be hindered by the complexities of homology determination in polyploid taxa. We use simulations to show that phasing alleles from allopolyploid individuals can improve phylogenetic network inference under the multispecies coalescent by obtaining the true network with fewer loci compared with haplotype consensus sequences or sequences with heterozygous bases represented as ambiguity codes. Phased allelic data can also improve divergence time estimates for networks, which is helpful for evaluating allopolyploid speciation hypotheses and proposing mechanisms of speciation. To achieve these outcomes in empirical data, we present a novel pipeline that leverages a recently developed phasing algorithm to reliably phase alleles from polyploids. This pipeline is especially appropriate for target enrichment data, where the depth of coverage is typically high enough to phase entire loci. We provide an empirical example in the North American Dryopteris fern complex that demonstrates insights from phased data as well as the challenges of network inference. We establish that our pipeline (PATÉ: Phased Alleles from Target Enrichment data) is capable of recovering a high proportion of phased loci from both diploids and polyploids. These data may improve network estimates compared with using haplotype consensus assemblies by accurately inferring the direction of gene flow, but statistical nonidentifiability of phylogenetic networks poses a barrier to inferring the evolutionary history of reticulate complexes.

Evolutionary Biology↗

Duke Forest FACE (FACTS-I): Meteorological and Soil Data

This dataset, collected from Duke Forest Free Air CO2 Enrichment (FACE) – Forest-Atmosphere Carbon Transfer and Storage (FACTS-I) experiment, includes variables describing the meteorological conditions above canopy, within canopy, and soil depending on the variable. The Duke FACE experiment was located in a loblolly pine (Pinus taeda L.) plantation established in 1983. Naturally regenerated broadleaved species including sweetgum (Liquidambar styraciflua L.) and tulip poplar (Liriodendron tulipifera L.), mostly in the overstory, and winged elm (Ulmus alata Michx.) and red maple (Acer rubrum L.) were common in the understory. The FACE experiment commenced with two plots (plots 7-8) in 1994 (Oren et al. 2001), with six additional plots (plots 1-6) coming online on 27 August 1996. The CO2 enrichment was terminated on 31 October 2010 and post-enrichment data collection continued through 2012. Complete fertilization was applied annually to half of plots 7-8 from 1998 to 2004. The nutrient addition experiment expanded to half of plot 1-6 with a common protocol of N-fertilization in 2005. N-fertilization continued until 2012. The data range varied by sensor availability. A summary of information about variable name and data range can be found in the ‘FileDescription_[variable_name].txt’ files.

54 ENVIRONMENTAL SCIENCES↗

International Space Station (ISS) Anomalies Trending Study: Appendices - Volume II

The NASA Engineering and Safety Center (NESC) set out to utilize data mining and trending techniques to review the anomaly history of the International Space Station (ISS) and provide tools for discipline experts not involved with the ISS Program to search anomaly data to aid in identification of areas that may warrant further investigation. Additionally, the assessment team aimed to develop an approach and skillset for integrating data sets, with the intent of providing an enriched data set for discipline experts to investigate that is easier to navigate, particularly in light of ISS aging and the plan to extend its life into the late 2020s. This document contains the Appendices to the Volume I report.

Beil, Robert J.↗

International Space Station (ISS) Anomalies Trending Study

The NASA Engineering and Safety Center (NESC) set out to utilize data mining and trending techniques to review the anomaly history of the International Space Station (ISS) and provide tools for discipline experts not involved with the ISS Program to search anomaly data to aid in identification of areas that may warrant further investigation. Additionally, the assessment team aimed to develop an approach and skillset for integrating data sets, with the intent of providing an enriched data set for discipline experts to investigate that is easier to navigate, particularly in light of ISS aging and the plan to extend its life into the late 2020s. This report contains the outcome of the NESC Assessment.

Beil, Robert J.↗

Demonstration of neutrinoless double beta decay searches in gaseous xenon with NEXT

The NEXT experiment aims at the sensitive search of the neutrinoless double beta decay in 136 Xe, using high-pressure gas electroluminescent time projection chambers. The NEXT-White detector is the first radiopure demonstrator of this technology, operated in the Laboratorio Subterráneo de Canfranc. Achieving an energy resolution of 1% FWHM at 2.6 MeV and further background rejection by means of the topology of the reconstructed tracks, NEXT-White has been exploited beyond its original goals in order to perform a neu- trinoless double beta decay search. The analysis considers the combination of 271.6 days of 136 Xe-enriched data and 208.9 days of 136Xe-depleted data. A detailed background modeling and measurement has been developed, ensuring the time stability of the radiogenic and cosmogenic contributions across both data samples. Limits to the neutrinoless mode are obtained in two alternative analyses: a background-model-dependent approach and a novel direct background-subtraction technique, offering results with small dependence on the background model assumptions. With a fiducial mass of only 3.50 ± 0.01 kg of 136 Xe-enriched xenon, 90% C.L. lower limits to the neutrinoless double beta decay are found in the $T^{0v}_{1/2} > 5.5 \times 10^{23} - 1.3 \times 10^{24}$ yr range, depending on the method. The presented techniques stand as a proof-of-concept for the searches to be implemented with larger NEXT detectors.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Measurement of inclusive and differential cross sections for W + W − production in proton-proton collisions at $\sqrt{s} = 13$ TeV

Measurements at $\sqrt{s} = 13$ TeV of the opposite-sign W boson pair production cross section in proton-proton collisions are presented. The data used in this study were collected with the CMS detector at the CERN LHC in 2022, and correspond to an integrated luminosity of 34.8 fb -1 . Events are selected by requiring one electron and one muon of opposite charge. A maximum likelihood fit is performed on signal- and background-enriched data categories defined by the flavor and charge of the leptons, the number of jets, and number of jets originating from b quarks. The overall sensitivity is significantly better than that of previous results with a similar integrated luminosity. The improvement comes from a more refined control of experimental uncertainties and an improved fit strategy. An inclusive W + W - production cross section of 125.7 ± 5.6 pb is measured, in agreement with standard model predictions. Cross sections are also reported in a fiducial region close to that of the detector acceptance, both inclusively and differentially, as a function of the jet multiplicity in the event. For the first time in proton-proton collisions, WW events with zero, one, and at least two jets are studied simultaneously and compared with recent theoretical predictions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Developing Asparagaceae1726: An Asparagaceae‐specific probe set targeting 1726 loci for Hyb‐Seq and phylogenomics in the family

Abstract Premise Target sequence capture (Hyb‐Seq) is a cost‐effective sequencing strategy that employs RNA probes to enrich for specific genomic sequences. By targeting conserved low‐copy orthologs, Hyb‐Seq enables efficient phylogenomic investigations. Here, we present Asparagaceae1726—a Hyb‐Seq probe set targeting 1726 low‐copy nuclear genes for phylogenomics in the angiosperm family Asparagaceae—which will aid the often‐challenging delineation and resolution of evolutionary relationships within Asparagaceae. Methods Here we describe and validate the Asparagaceae1726 probe set (https://github.com/bentzpc/Asparagaceae1726) in six of the seven subfamilies of Asparagaceae. We perform phylogenomic analyses with these 1726 loci and evaluate how inclusion of paralogs and bycatch plastome sequences can enhance phylogenomic inference with target‐enriched data sets. Results We recovered at least 82% of target orthologs from all sampled taxa, and phylogenomic analyses resulted in strong support for all subfamilial relationships. Additionally, topology and branch support were congruent between analyses with and without inclusion of target paralogs, suggesting that paralogs had limited effect on phylogenomic inference. Discussion Asparagaceae1726 is effective across the family and enables the generation of robust data sets for phylogenomics of any Asparagaceae taxon. Asparagaceae1726 establishes a standardized set of loci for phylogenomic analysis in Asparagaceae, which we hope will be widely used for extensible and reproducible investigations of diversification in the family.

Plant Sciences↗

Derivation of physical equations for high-speed laser welding using large language models

It is challenging to formulate complex physical phenomena that occur in a manufacturing process, particularly when the available data are limited, rendering conventional data-driven approaches ineffective. This study aims to predict humping onset in high-speed laser welding by introducing a novel framework, namely text-to-equations generative pre-trained transformer (T2EGPT). This method leverages the capabilities of large language models (LLMs), in combination with sparse experimental data and enriched literature data, to derive an interpretable and generalizable equation for predicting humping initiation. By capturing key correlations among physical parameters, T2EGPT generates a compact and dimensionless expression that accurately predicts hump formation. The equation reveals that humping arises from the interplay between inertia-driven backward melt flow and capillary-driven surface stabilization, where inertial forces drive molten metal backward and capillary forces resist surface deformation. Furthermore, compared to traditional data-driven models, T2EGPT demonstrates enhanced predictive accuracy and cross-material transferability. More broadly, this study highlights the potential of LLMs to integrate textual information with data-driven discovery, enabling the extraction of physical laws in data-scarce scientific domains.

36 MATERIALS SCIENCE↗

Autoregressive long-horizon prediction of plasma edge dynamics *

Accurate modeling of scrape-off layer (SOL) and divertor-edge dynamics is vital for designing plasma-facing components in fusion devices. High-fidelity edge fluid/neutral codes such as SOLPS-ITER capture SOL physics with high accuracy, but their computational cost limits broad parameter scans and long transient studies. We present transformer-based, autoregressive surrogates for efficient prediction of 2D, time-dependent plasma edge state fields. Trained on SOLPS-ITER spatiotemporal data for the KSTAR tokamak, the surrogates forecast electron temperature, electron density, and radiated power over extended horizons. We evaluate model variants trained with increasing autoregressive horizons (1–100 steps) on short- and long-horizon prediction tasks. Longer-horizon training systematically improves rollout stability and mitigates error accumulation, enabling stable predictions over hundreds to thousands of steps and reproducing key dynamical features such as the motion of high-radiation regions. Measured end-to-end wall-clock times show the surrogate is orders of magnitude faster than SOLPS-ITER, enabling rapid parameter exploration. Prediction accuracy degrades when the surrogate enters physical regimes not represented in the training dataset, motivating future work on data enrichment and physics-informed constraints. Overall, this approach provides a fast, accurate surrogate for computationally intensive plasma edge simulations, supporting rapid scenario exploration, control-oriented studies, and progress toward real-time applications in fusion devices.

autoregressive deep learning↗

Software Bill of Materials (SBOM) Sharing Lifecycle Report

As Software Bill of Materials (SBOM) adoption efforts mature, SBOM sharing continues to occur, but no single solution or set of solutions have become ubiquitous. The purpose of this report is to enumerate and describe the different parties and phases of the SBOM sharing lifecycle and assist readers in choosing suitable SBOM sharing solutions based on the amount of time, resources, subject-matter expertise, effort, and access to tooling that is available to the reader to implement a phase of the SBOM sharing lifecycle. The SBOM sharing lifecycle consists of the Discovery, Access, and Transport of an SBOM and this report details these individual phases and how an SBOM goes from author to the consumer. This report also details how potential enrichment activities may be performed on an SBOM to create a new product before or after it has been shared. The concept of a sophistication classification for SBOM sharing solutions is concurrently introduced with a focus on the inclusion or lack of certain features and effort associated with their implementation. Examples of low, medium, and high-sophistication solutions are provided; however, these examples and associated categorizations should not be seen as a qualitative judgment meant to push the reader towards a particular adoption strategy since sharing solutions are chosen based on the unique needs of the user. This report does recommend the SBOM community consider how to make current and future sharing solutions interoperable with each other as well as more automated methods to facilitate sharing and broader SBOM adoption. This report also highlights an SBOM sharing survey results obtained from interviews with stakeholders to understand the current SBOM sharing landscape. The categorized results of the survey suggest that SBOMs are currently transported directly to the receiver through email or similar informal communication mechanisms or alternatively the SBOM resides on a repository available to consumers. In addition to these transport methods, this report captures industry efforts to create private sharing solutions and services that can store and transport enrichment data and may use higher sophistication features that are cloud-based or using distributed ledger technologies.

97 MATHEMATICS AND COMPUTING↗

Some data processing requirements for precision Nap-Of-the-Earth (NOE) guidance and control of rotorcraft

Nap-Of-the-Earth (NOE) flight in a conventional helicopter is extremely taxing for two pilots under visual conditions. Developing a single pilot all-weather NOE capability will require a fully automatic NOE navigation and flight control capability for which innovative guidance and control concepts were examined. Constrained time-optimality provides a validated criterion for automatically controlled NOE maneuvers if the pilot is to have confidence in the automated maneuvering technique. A second focus was to organize the storage and real-time updating of NOE terrain profiles and obstacles in course-oriented coordinates indexed to the mission flight plan. A method is presented for using pre-flight geodetic parameter identification to establish guidance commands for planned flight profiles and alternates. A method is then suggested for interpolating this guidance command information with the aid of forward and side looking sensors within the resolution of the stored data base, enriching the data content with real-time display, guidance, and control purposes. A third focus defined a class of automatic anticipative guidance algorithms and necessary data preview requirements to follow the vertical, lateral, and longitudinal guidance commands dictated by the updated flight profiles and to address the effects of processing delays in digital guidance and control system candidates. The results of this three-fold research effort offer promising alternatives designed to gain pilot acceptance for automatic guidance and control of rotorcraft in NOE operations.

Clement, Warren F.↗

Interpreting omics data with pathway enrichment analysis

Pathway enrichment analysis is indispensable for interpreting omics datasets and generating hypotheses. However, the foundations of enrichment analysis remain elusive to many biologists. Here, in this study, we discuss best practices in interpreting different types of omics data using pathway enrichment analysis and highlight the importance of considering intrinsic features of various types of omics data. We further explain major components that influence the outcomes of a pathway enrichment analysis, including defining background sets and choosing reference annotation databases. To improve reproducibility, we describe how to standardize reporting methodological details in publications. This article aims to serve as a primer for biologists to leverage the wealth of omics resources and motivate bioinformatics tool developers to enhance the power of pathway enrichment analysis.

60 APPLIED LIFE SCIENCES↗

Enriching OpenStreetMap network data for transportation applications: Insights into the impact of urban congestion on accessibility

OpenStreetMap (OSM) data is a valuable open-source resource for various transportation, traffic, and planning applications. However, OSM network data lack operating traffic speed information, which is critical for transport planning and operations. Addressing this shortcoming, this study leverages commercial vendor data (to serve as ground truth) with exogenous, open-source variables characterizing local transport infrastructure, land use, and demographic information to predict average congested traffic speeds on OSM networks. Three machine-learning models were tested and estimated for OSM links with and without speed limit information in the Denver metropolitan region. Among these, XGBoost performed best, with mean absolute errors of 3.27 and 3.62 mph for links with and without speed limits, respectively. The developed models accurately predicted traffic speeds for different hours and days of the week compared to ground truth data. Using these predicted speeds, drive accessibility scores were computed for the Denver region for different time periods using the Mobility Energy Productivity (MEP) metric to understand the impact of congestion on energy-efficient accessibility. Results show that congestion-adjusted drive accessibility can be significantly lower compared to accessibility calculated using free flow speeds. Specifically, weekday evening hours saw a 42 % drop in accessibility due to reduced speeds, particularly around downtown Denver. Across the Denver metro region, approximately half as many opportunities and jobs are accessible in under 20 min by car during the evening peak period relative to free flow conditions. These findings underscore the importance of using congestion-adjusted operating speeds rather than speed limits in accessibility calculations, as reliance on speed limits can substantially overestimate energy-efficient drive accessibility in large, car-centric cities susceptible to significant congestion. In conclusion, the methodology presented here could further enrich OSM network data, making them useful for an even broader range of transportation applications.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Generation of Enrichment-Dependent Thermal Neutron Scattering Data

This work details the generation of enrichment-dependent thermal neutron scattering cross sections for several crucial uranium fuel compounds. The evaluations of the thermal scattering law (TSL) and associated cross sections for uranium dioxide (UO 2 ), uranium carbide (UC), and uranium nitride (UN) were performed using standard ab initio lattice dynamics (AILD) methods. The data for uranium metal was produced using a novel hybrid approach of molecular dynamics combined with lattice dynamics methods. 235 U enrichments of 5%, 10% (LEU+), 19.75% (HALEU), 93% (HEU), and 100% were considered, in addition to natural uranium. The enrichment-dependent masses and free atom cross sections were used in the generation of elastic and inelastic thermal neutron scattering cross sections, while the calculation of the phonon density of states (DOS) and resulting TSL considered only the natural isotopic composition of uranium. The use of an identical DOS for all enrichments is expected to have minimal impact on the final data, as the small change in uranium mass should not significantly affect lattice vibrations. The cross sections are shown to exhibit significant dependence on 235 U enrichment. The submission of this data to the National Nuclear Data Center (NNDC) for release in the ENDF/B-VIII.1 database should support the design of advanced reactor concepts.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗