Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Energy-efficient cooperative resource allocation and task scheduling for Internet of Things environments

Offloading Internet of Things (IoT) tasks to the cloud for further processing might not always lead to an optimal execution time, particularly in situations such as resource contention, under-provisioning, over-provisioning, and fragmentation. In addition, dynamically optimizing the number of Virtual Machines (VMs) for resource scheduling in order to meet application requirements remains a major research challenge. Further, existing resource scheduling algorithms focus primarily on minimizing operational costs while maximizing resource sharing and utilization. Considering energy utilization as part of the resource allocation and scheduling process as an optimization objective for maintaining load balancing has often been neglected. To address these challenges and more, we propose a cooperative energy-aware resource allocation and scheduling strategy based on a Technique for Order of Preference by Similarity to Ideal Solution (TOPSIS) multi-criteria decision-making method. Here we used the Grid Workloads Archive dataset to evaluate our proposed approach named TOPREAL. Experimental results with respect to the allocation of VM resources when considering processing a large segment of tasks indicate that TOPREAL outperforms existing algorithms in terms of energy savings, with an average improvement of 40.25%, while maintaining an average improvement of 16.21% when it comes to execution time. Results also demonstrate that our method can save an average of 78.06 processing hours and 63,215kJ of energy when compared to existing scheduling algorithms. These results demonstrate the effectiveness of our proposed model and the viability of using multi-criteria decision-making techniques such as TOPSIS to solve the resource allocation and scheduling problem in edge environments.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

FL‐ADS: Federated learning anomaly detection system for distributed energy resource networks

Abstract With the ongoing development of Distributed Energy Resources (DER) communication networks, the imperative for strong cybersecurity and data privacy safeguards is increasingly evident. DER networks, which rely on protocols such as Distributed Network Protocol 3 and Modbus, are susceptible to cyberattacks such as data integrity breaches and denial of service due to their inherent security vulnerabilities. This paper introduces an innovative Federated Learning (FL)‐based anomaly detection system designed to enhance the security of DER networks while preserving data privacy. Our models leverage Vertical and Horizontal Federated Learning to enable collaborative learning while preserving data privacy, exchanging only non‐sensitive information, such as model parameters, and maintaining the privacy of DER clients' raw data. The effectiveness of the models is demonstrated through its evaluation on datasets representative of real‐world DER scenarios, showcasing significant improvements in accuracy and F1‐score across all clients compared to the traditional baseline model. Additionally, this work demonstrates a consistent reduction in loss function over multiple FL rounds, further validating its efficacy and offering a robust solution that balances effective anomaly detection with stringent data privacy needs.

Purohit, Shaurya [Iowa State University Ames Iowa ↗

Fast correlation function calculator: A high-performance pair-counting toolkit

A novel high-performance exact pair-counting toolkit called fast correlation function calculator (FCFC) is presented. With the rapid growth of modern cosmological datasets, the evaluation of correlation functions with observational and simulation catalogues has become a challenge. High-efficiency pair-counting codes are thus in great demand. We introduce different data structures and algorithms that can be used for pair-counting problems, and perform comprehensive benchmarks to identify the most efficient algorithms for real-world cosmological applications. We then describe the three levels of parallelisms used by FCFC, SIMD, OpenMP, and MPI, and run extensive tests to investigate the scalabilities. Finally, we compare the efficiency of FCFC with alternative pair-counting codes. The data structures and histogram update algorithms implemented in FCFC are shown to outperform alternative methods. FCFC does not benefit greatly from SIMD because the bottleneck of our histogram update algorithm is mainly cache latency. Nevertheless, the efficiency of FCFC scales well with the numbers of OpenMP threads and MPI processes, even though speedups may be degraded with over a few thousand threads in total. FCFC is found to be faster than most (if not all) other public pair-counting codes for modern cosmological pair-counting applications.

79 ASTRONOMY AND ASTROPHYSICS↗

Emerging Flexible Designs for Geospatial Multimodal Foundation Models

Foundation models are rapidly transforming Earth observation by enabling scalable pretraining across diverse unlabeled geospatial modalities. However, their architectural diversity—ranging from encoder-only to encoder-decoder and masked autoencoding paradigms—makes it challenging to assess performance trade-offs in a consistent manner. In this work, we present an apples-to-apples comparison of leading FM architectures designed for geospatial multimodal reasoning, with a particular focus on flexibility across varied spectral band configurations. We standardize pretraining using identical self-supervised learning objectives and training datasets, and evaluate all models under consistent parameterization on the GEOBench benchmark across classification and segmentation tasks. Our results offer new insights into the design trade-offs between model flexibility, modality alignment, and downstream task performance. By highlighting architectural strengths and limitations under controlled conditions, this study provides practical guidance for building next-generation geospatial foundation models capable of robust multimodal reasoning.

Ambrozio Dias, Philipe [ORNL] (ORCID:0000000194277↗

Grid Optimization (GO) Competition Platform

A software for a multi-challenge power-flow grid optimization competition was developed. The platform brings together high performance computing clusters, webservers, databases, competition datasets, schedulers, evaluation codes, and a multitude of language compilers and optimization solvers to host the competition. The original video announcing the competition, from former Secretary Perry, at: https://www.youtube.com/watch?v=hZwX3P9vS8M

Veeramany, Arun↗

FATHOMS-RAG

Dataset and Evaluation code for the FATHOMS-RAG paper

Hildebrand, Samuel [Louisiana State Univ., Baton R↗

HITMAN

HITMAN (Hermite Interpolation of Trajectories and Measurement Synthesis for Analysis of Navigators) interpolates—or estimates the unknown values between known values—flight trajectories and generates synthetic inertial measurement unit (IMU) data using Hermite splines. This Python library provides modeling and simulation capabilities to synthesize inertial measurements from discrete trajectory points, enabling researchers to create exemplar datasets for evaluating navigation algorithms in various applications, including consumer devices like smartphones and vehicles. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Walker II, Michael [Sandia National Lab. (SNL-CA),↗

TULIP: An RNA-seq-based Primary Tumor Type Prediction Tool Using Convolutional Neural Networks

Background: With cancer as one of the leading causes of death worldwide, accurate primary tumor type prediction is critical in identifying genetic factors that can inhibit or slow tumor progression. There have been efforts to categorize primary tumor types with gene expression data using machine learning, and more recently with deep learning, in the last several years. Methods In this paper, we developed four 1-dimensional (1D) Convolutional Neural Network (CNN) models to classify RNA-seq count data as one of 17 highly represented primary tumor types or 32 primary tumor types regardless of imbalanced representation. Additionally, we adapted the models to take as input either all Ensembl genes (60,483) or protein coding genes only (19,758). Unlike previous work, we avoided selection bias by not filtering genes based on expression values. RNA-seq count data expressed as FPKM-UQ of 9,025 and 10,940 samples from The Cancer Genome Atlas (TCGA) were downloaded from the Genomic Data Commons (GDC) corresponding to 17 and 32 primary tumor types respectively for training and validating the models. Results: All 4 1D-CNN models had an overall accuracy of 94.7% to 97.6% on the test dataset. Further evaluation indicates that the models with protein coding genes only as features performed with better accuracy compared to the models with all Ensembl genes for both 17 and 32 primary tumor types. For all models, the accuracy by primary tumor type was above 80% for most primary tumor types. Conclusions: We packaged all 4 models as a Python-based deep learning classification tool called TULIP (TUmor CLassIfication Predictor) for performing quality control on primary tumor samples and characterizing cancer samples of unknown tumor type. Further optimization of the models is needed to improve the accuracy of certain primary tumor types.

Jones, Sara↗

Data-Driven Energy Resilience Assessment and Enhancement in Urban Communities: A Case Study in Detroit

This paper presents a data-driven framework for assessing and enhancing energy resilience in urban communities. The resilience assessment is based on two datasets: 1) annual aggregated power outage data and 2) 15-minute interval outage data. High-impact, low-probability (HILP) events are identified within these datasets to evaluate community resilience under extreme conditions. To enhance resilience, an optimization framework utilizing mixed integer linear programming is developed to determine the optimal sizing and placement of solar photovoltaic (PV) systems and battery energy storage systems (BESS). This method offers a cost-effective and practical solution for improving energy resilience in vulnerable communities. Furthermore, a case study of the City of Detroit in Michigan demonstrates the effectiveness of the framework through simulation and validation.

Energy resilience assessment↗

Dispersion Validation for Flow Involving a Large Structure Revisited: 45 Degree Rotation

The atmospheric dispersion of contaminants in the wake of a large urban structure is a challenging fluid mechanics problem of interest to the scientific and engineering communities. Magnetic Resonance Velocimetry (MRV) and Magnetic Resonance Concentration (MRC) are relatively new techniques that leverage diagnostic equipment used primarily by the medical field to make 3D engineering measurements of flow and contaminant dispersal. SIERRA/Fuego, a computational fluid dynamics (CFD) code at Sandia National Labs is employed to make detailed comparisons to the dataset to evaluate the quantitative and qualitative accuracy of the model. This work is the second in a series of scenarios. In the prior work, a single large building in an array of similar buildings was considered with the wind perpendicular to a building face. In this work, the geometry is rotated by 45 degrees and improved studies are performed for simulation credibility. The comparison exercise shows conditionally good comparisons between the model and experiment. Model uncertainties are assessed through parametric variations. Various methods of quantifying the accuracy between experiments and data are examined Three-dimensional analysis of accuracy is performed. The effort helped identify deficiencies in the techniques used to make these comparisons, and further methods development therefore becomes one of the main recommendations for follow-on work.

42 ENGINEERING↗

Fuel-Cladding Eutectic Study of Legacy Fast Flux Test Facility (FFTF) MFF HT9/U-10Zr Metallic Fuel

This report presents the first systematic investigation of fuel-cladding eutectic interaction (FCEI) in irradiated HT9/U-10Zr metallic fuel from the Fast Flux Test Facility (FFTF) Materials Fuels Form (MFF) program, using differential scanning calorimetry (DSC) coupled with scanning electron microscopy (SEM) and energy dispersive X-ray spectroscopy (EDS). Two irradiated fuel cross-sections, MNT07H (9.5 at% burnup, x/L = 0.78) and MNT08H (7.0 at% burnup, x/L = 0.93), were subjected to three successive isothermal annealing rounds (R1–R3) at 820°C for 20 minutes each, yielding a cumulative transient duration of one hour. This study directly addresses a recognized gap in the existing FCEI database, which previously lacked irradiated HT9/U-10Zr data at burnup levels above 8 at%. Two principal findings emerge from the study. First, for both samples, FCEI remained spatially confined within the pre-existing fuel-cladding chemical interaction (FCCI) zone boundaries after R3, with no measurable eutectic penetration into unaffected cladding beyond the original FCCI layer. This self-limiting behavior is consistent with historical Fuel Behavior Test Apparatus (FBTA) results and is attributed to the near-eutectic phase composition of the FCCI zone, which rapidly absorb the available eutectic-forming constituents and then stall penetration once the FCCI zone is consumed. A comparison with unirradiated surrogate data further supports this mechanism: whereas a U–34 at.% Fe sample would be expected to show ~176 µm of iron penetration under comparable conditions, the irradiated samples exhibited only ~20 µm, a discrepancy attributed to irradiation-induced interfacial porosity and pre-existing FCCI composition gradients. Second, for MNT08H, FCEI was observed exclusively on the half of the cladding circumference where pre-existing steady-state FCCI was present, with no detectable FCEI on the opposite half. Three hypotheses are proposed to explain this asymmetry: the inhibiting role of a zirconium-rich rind at the fuel-cladding interface; the chemical sequestration of iron by redistributed zirconium within the fuel matrix; and the persistence of fuel-cladding gaps on the FCEI-free half that preclude direct contact. All three hypotheses require further experimental investigation. The results extend the empirical FCEI database into higher-burnup territory and demonstrate the viability of DSC-based testing as a substitute for the no-longer-available FBTA apparatus. Future work will include additional cross-section testing, compilation of the full FCEI dataset, model evaluation, and DSC testing of ternary fuel compositions to broaden the experimental basis for safety assessment of sodium-cooled fast reactor systems.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Real-Time Anomaly Detection for Beyond Standard Model Searches in ProtoDUNE Horizontal Drift

This paper summarizes work conducted throughout a SULI internship at Fermi National Accelerator Laboratory focused on building an unsupervised machine learning model for real-time anomaly detection in ProtoDUNE Horizontal Drift. Using simulated data, we trained an autoencoder model on a pure cosmic dataset, and evaluated it on both cosmic and neutrino events---making the model an anomaly detector. The goal was to make a model which matches or exceeds the current ADC Simple Window trigger algorithm so that our model can perform at the same rate but provide sensitivity to potential beyond-the-Standard-Model (BSM) signatures. In the end, we were able to construct a model which slightly exceeds the capabilities of the ADC Simple Window while remaining completely unsupervised, achieving $31.9 \pm 0.2$\% ($26.6 \pm 0.2$\%) $\nu$ efficiency at 5 Hz (2 Hz), a 3.6 (3.2) percentage point increase. Additionally, $17.5 \pm 0.3$\% ($18.3 \pm 0.3$\%) of the events that passed the autoencoder at 5 Hz (2 Hz) were missed by the current trigger algorithm. Future work will investigate alternative normalization methods, including quantile transformation, and evaluate the model on ProtoDUNE-HD detector-glitch data if that data becomes available.

Wilson, Cameron C. [Cincinnati U., RWC]↗

Real-Time Anomaly Detection for Beyond Standard Model Searches in ProtoDUNE Horizontal Drift

This paper summarizes work conducted throughout a SULI internship at Fermi National Accelerator Laboratory focused on building an unsupervised machine learning model for real-time anomaly detection in ProtoDUNE Horizontal Drift. Using simulated data, we trained an autoencoder model on a pure cosmic dataset, and evaluated it on both cosmic and neutrino events---making the model an anomaly detector. The goal was to make a model which matches or exceeds the current ADC Simple Window trigger algorithm so that our model can perform at the same rate but provide sensitivity to potential beyond-the-Standard-Model (BSM) signatures. In the end, we were able to construct a model which slightly exceeds the capabilities of the ADC Simple Window while remaining completely unsupervised, achieving $31.9 \pm 0.2$\% ($26.6 \pm 0.2$\%) $\nu$ efficiency at 5 Hz (2 Hz), a 3.6 (3.2) percentage point increase. Additionally, $17.5 \pm 0.3$\% ($18.3 \pm 0.3$\%) of the events that passed the autoencoder at 5 Hz (2 Hz) were missed by the current trigger algorithm. Future work will investigate alternative normalization methods, including quantile transformation, and evaluate the model on ProtoDUNE-HD detector-glitch data if that data becomes available.

Wilson, Cameron C. [Cincinnati U., RWC]↗

Fire and Drought Affect Multiple Aspects of Diversity in a Migratory Bird Stopover Community

Drought and high-severity, stand-replacing wildfires can have substantial impacts on the composition of avian communities, including stop-over communities during migration. An inextricable link exists between drought and wildfire, each operating and impacting across different timescales. Many studies have found nonlinear avian abundance trends in breeding community time series data that include pre- and post-fire observations, describing an initial decrease in abundance followed by rapid increases that can attenuate over time. Here, we use a fall bird-banding dataset to evaluate shifts in a drought-impacted avian community following wildfire from taxonomic, functional, and phylogenetic perspectives. We looked at the community as a whole and also categorized birds as residents, migrants, and breeders to assess potential varying responses at the study site. We observed post-fire shifts in functional and phylogenetic diversity that corresponded to changes in vegetation. An influx of migratory insectivores post-fire drove much of the variation between pre- and post-fire avian communities and toward a more related, less phylogenetically dispersed community. A concurrent monsoon season drought was also associated with functional and phylogenetic diversity, highlighting the intertwined pulse press effects on avian communities. Overall, our results suggest that, although bird communities are immediately impacted by fire-driven resource changes, they can rebound over time, it is unclear how long-term drought may continue to shape the composition of these avian communities.

59 BASIC BIOLOGICAL SCIENCES↗

Assessment of Climate Change Impacts on Renewable Energy Resources in Western North America

We examine a 25 km resolution climate model dataset to evaluate how regional climate change impacts solar and wind energy under a high-emission scenario. Our study considers the Western Electricity Coordinating Council (WECC) region, which covers the western United States and southwestern Canada, focusing specifically on locations with existing solar and wind infrastructure. First, we conduct a historical model comparison of solar and wind energy capacity factors to highlight model uncertainties across the study area. Using future climate projections, we then assess the seasonal patterns of solar and wind capacity factors for three timeframes: historical, mid-century, and end of century. Additionally, we estimate the frequency of solar and wind resource droughts during these periods for the entire WECC and its five operational subregions, finding that certain subregions are more susceptible to energy droughts due to limited renewable resources. Finally, we present day-ahead capacity factor forecasts to support energy storage planning and provide estimates of offshore wind energy capacity within the WECC. Our results indicate that offshore wind capacity factors are nearly twice as high as onshore values, with less seasonal variation, which suggests that offshore wind could offer a more consistent renewable energy supply in the future.

climate change↗

Atmospheric Radiation Measurement (ARM) airborne field campaign data products between 2013 and 2018

Airborne measurements are pivotal for providing detailed, spatiotemporally resolved information about atmospheric parameters and aerosol and cloud properties, thereby enhancing our understanding of dynamic atmospheric processes. For 30 years, the US Department of Energy (DOE) Office of Science supported an instrumented Gulfstream 1 (G-1) aircraft for atmospheric field campaigns. Data from the final decade of G-1 operations were archived by the Atmospheric Radiation Measurement (ARM) Data Center and made publicly available at no cost to all registered users. To ensure a consistent data format and to improve the accessibility of the ARM airborne data, an integrated dataset was recently developed covering the final 6 years of G-1 operations (2013 to 2018, https://doi.org/10.5439/1999133; Mei and Gaustad, 2024). The integrated dataset includes data collected from 236 flights (766.4 h), which covered the Arctic, the US Southern Great Plains (SGP), the US West Coast, the eastern North Atlantic (ENA), the Amazon Basin in Brazil, and the Sierras de Córdoba range in Argentina. These comprehensive data streams provide much-needed insight into spatiotemporal variability in the thermodynamic quantities and aerosol and cloud properties for addressing essential science questions in Earth system process studies. This paper describes the DOE ARM merged G-1 datasets, including information on the acquisition, data collection challenges and future potentials, and quality control processes. It further illustrates the usage of this merged dataset to evaluate the Energy Exascale Earth System Model (E3SM) with the Earth System Model Aerosol–Cloud Diagnostics (ESMAC Diags) package.

54 ENVIRONMENTAL SCIENCES↗

Real-Time Anomaly Detection for Searches Beyond the Standard Model in the ProtoDUNE Horizontal Drift Detector

This paper summarizes work conducted throughout a SULI internship at Fermi National Accelerator Laboratory focused on building an unsupervised machine learning model for real-time anomaly detection in ProtoDUNE Horizontal Drift. Using simulated data, we trained an autoencoder model on a pure cosmic dataset, and evaluated it on both cosmic and neutrino events—making the model an anomaly detector. The goal was to make a model which matches or exceeds the current ADC Simple Window trigger algorithm so that our model can perform at the same rate but provide sensitivity to potential beyond-the-Standard-Model (BSM) signatures. In the end, we were able to construct a model which slightly exceeds the capabilities of the ADC Simple Window while remaining completely unsupervised, achieving 31.9 ± 0.2% (26.6 ± 0.2%) ν efficiency at 5 Hz (2 Hz), a 3.6 (3.2) percentage point increase. Additionally, 17.5 ± 0.3% (18.3 ± 0.3%) of the events that passed the autoencoder at 5 Hz (2 Hz) were missed by the current trigger algorithm. Future work will investigate alternative normalization methods, including quantile transformation, and evaluate the model on ProtoDUNE-HD detector-glitch data if that data becomes available.

Wilson, C. [Cincinnati U., RWC]↗