Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A 1 km soil moisture dataset over eastern CONUS generated by assimilating SMAP data into the Noah-MP land surface model

An improved fine-scale soil moisture (SM) dataset at 1 km grid spacing, covering much of the eastern continental US, was generated by assimilating 9 km Soil Moisture Active Passive (SMAP) SM data into the v4.0.1 Noah-MP land surface model. With 12 ensemble members, the assimilation was carried out using the ensemble Kalman filter algorithm within NASA's Land Information System. The SM analysis for 2016 was fully validated against in situ observations from four different networks and compared with four other existing datasets. Results indicate that this SM analysis surpasses other datasets in top-layer SM distribution, including a machine-learning-based product, despite all SM estimates being less heterogeneous than observed. The analysis of anomalous errors suggests that large similarity in intrinsic errors is likely due to overlapping data sources among the selected SM datasets. More detailed evaluations were performed over two geographic areas. The observations collected by the Atmospheric Radiation Measurement facility in Oklahoma suggest that soil temperature and surface heat fluxes are concurrently simulated with good accuracy. Investigation into the 2016 southeastern US drought response further indicates drier conditions and higher evapotranspiration estimates compared to GLEAMv4.1. Notably, large errors are associated with grids having clay soil textures, underscoring the need for refined model treatments for specific soil types to further improve SM estimates. The dataset is publicly available on Zenodo at https://doi.org/10.5281/zenodo.14370563 (Tai et al., 2024).

Tai, Sheng-Lun [Pacific Northwest National Laborat↗

Shortwave Array Spectroradiometer-Hemispheric (SAS-He): design and evaluation

A novel ground-based radiometer, referred to as the Shortwave Array Spectroradiometer-Hemispheric (SAS-He), is introduced. This radiometer uses the shadow-band technique to report total irradiance and its direct and diffuse components frequently (every 30 s) with continuous spectral coverage (350–1700 nm) and moderate spectral (~2.5 nm ultraviolet–visible and ~6 nm shortwave-infrared) resolution. The SAS-He's performance is evaluated using integrated datasets collected over coastal regions during three field campaigns supported by the US Department of Energy's Atmospheric Radiation Measurement (ARM) program, namely the (1) Two-Column Aerosol Project (TCAP; Cape Cod, Massachusetts), (2) Tracking Aerosol Convection Interactions Experiment (TRACER; in and around Houston, Texas), and (3) Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE; La Jolla, California). We compare (i) aerosol optical depth (AOD) and total optical depth (TOD) derived from the direct irradiance, as well as (ii) the diffuse irradiance and direct-to-diffuse ratio (DDR) calculated from two components of the total irradiance. As part of the evaluation, both AOD and TOD derived from the SAS-He direct irradiance are compared to those provided by a collocated Cimel sunphotometer (CSPHOT) at five (380, 440, 500, 675, 870 nm) and two (1020, 1640 nm) wavelengths, respectively. Additionally, the SAS-He diffuse irradiance and DDR are contrasted with their counterparts offered by a collocated multifilter rotating shadowband radiometer (MFRSR) at six (415, 500, 615, 675, 870, 1625 nm) wavelengths. Overall, reasonable agreement is demonstrated between the compared products despite the challenging observational conditions associated with varying aerosol loadings and diverse types of aerosols and clouds. For example, the AOD- and TOD-related values of root mean square error remain within 0.021 at 380, 440, 500, 675, 870, 1020, and 1640 nm wavelengths during the three field campaigns.

47 OTHER INSTRUMENTATION↗

a priori uncertainty quantification of reacting turbulence closure models using Bayesian neural networks

While many physics-based closure model forms have been posited for the sub-filter scale (SFS) in large eddy simulation (LES), vast amounts of data available from direct numerical simulations (DNS) create opportunities to leverage data-driven modeling techniques. Albeit flexible, data-driven models still depend on the dataset and the functional form of the model chosen. Increased adoption of such models requires reliable uncertainty estimates both in the data-informed and out-of-distribution regimes. Here, in this work, we employ Bayesian neural networks (BNNs) to capture both epistemic and aleatoric uncertainties in a reacting flow model. In particular, we model the filtered progress variable scalar dissipation rate which plays a key role in the dynamics of turbulent premixed flames. We demonstrate that BNN models can provide unique insights about the structure of uncertainty of the data-driven closure models. We also propose a method for the incorporation of out-of-distribution information in a BNN, which can be used for out-of-distribution query detection. The efficacy of the model is demonstrated by a priori evaluation on a dataset consisting of a variety of flame conditions and fuels.

97 MATHEMATICS AND COMPUTING↗

Foundational Dataset for Developing Large-Sample Stream Temperature Models in the Conterminous United States

This dataset provides inputs, evaluation results, and trained weights from a large-sample Long Short-Term Memory (LSTM) model designed to predict daily stream temperatures across unregulated river reaches in the conterminous United States (CONUS). It includes dynamic meteorological and hydrologic forcings, static physiographic attributes, and model outputs from cross-validation experiments spanning 300 basins. It supports reproducible modeling, direct application for new basins, and provides data suitable for integration with reservoir and river simulations under current and future climates. It contains two .zip files described below · RQ-AI_runs.zip: Model outputs from 10-fold cross-validation experiments, including observed and predicted daily stream temperatures, along with test performance metrics for water years 2017–2019. Two versions are included: 1. Model trained and validated using subbasin-area weighted dynamic features. 2. Model trained and validated using whole-basin area weighted dynamic features. · RQ-AI_inputs.zip: Collection of all formatted dynamic and static predictor datasets (meteorological, hydrologic, and physiographic features) used in model training and analysis. Detailed instructions and data structure is held at the following GitLab repository: https://code.ornl.gov/tempwise/training.

Gomez-Velez, Jesus [Oak Ridge National Laboratory ↗

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio↗

Seasonal variations in composition and sources of atmospheric ultrafine particles in urban Beijing based on near-continuous measurements

Abstract. Understanding the composition and sources of atmospheric ultrafine particles (UFPs) is essential in evaluating their exposure risks. It requires long-term measurements with high time resolution, which are scarce to date. We performed near-continuous measurements of UFP composition during four seasons in urban Beijing using a thermal desorption chemical ionization mass spectrometer, accompanied by real-time size distribution measurements. We found that UFPs in urban Beijing are dominated by organic components, varying seasonally from 68 % to 81 %. CHO organics (i.e., molecules containing carbon, hydrogen, and oxygen) are the most abundant in summer, while sulfur-containing organics, some nitrogen-containing organics, nitrate, and chloride are the most abundant in winter. With the increase of particle diameter, the contribution of CHO organics decreases, while that of sulfur-containing and nitrogen-containing organics, nitrate, and chloride increases. Source apportionment analysis of the UFP organics indicates contributions from cooking and vehicle sources, photooxidation sources enriched in CHO organics, and aqueous/heterogeneous sources enriched in nitrogen- and sulfur-containing organics. The increased contributions of cooking, vehicle, and photooxidation components are usually accompanied by simultaneous increases in UFP number concentrations related to cooking emission, vehicle emission, and new particle formation, respectively, while the increased contribution of the aqueous/heterogeneous composition is usually accompanied by the growth of UFP mode diameters. The highest UFP number concentrations in winter are due to the strongest new particle formation, the strongest local primary particle number emissions, and the slowest condensational growth of UFPs to larger sizes. This study provides a comprehensive understanding of urban UFP composition and sources and offers valuable datasets for the evaluation of UFP exposure risks.

54 ENVIRONMENTAL SCIENCES↗

The Use of Thermal Cameras for Pedestrian Detection

Visible-range camera sensors have been widely used for pedestrian detection. However, most of the methods, which employ visible-range color cameras, do not perform well under low-light and no-light conditions, e.g. during night time. Since the working principle of thermal camera sensors is mainly based on temperature and not light, they have been employed for person detection to overcome the drawbacks of visible-range sensors under these conditions. Every object gives off thermal energy, which is captured by a thermal camera sensor. When an object becomes hotter, it emits more thermal energy, and is therefore captured as much brighter or vice versa. Yet, compared to visible-range cameras, there are many additional challenges that need to be addressed when detecting pedestrians from thermal camera images. These challenges include bright hot objects close to humans, similar pixel values in an image due to weather conditions, or objects that block thermal cameras such as concrete or glass. Glass acts like a mirror for infrared radiation and reflects whatever is in front of the camera. Thus, novel methods are still required to accomplish pedestrian detection task from thermal camera images. To contribute to these efforts, we propose a new method and a modified object detection network incorporating saliency maps of thermal camera images. The features obtained from thermal images and their corresponding saliency maps are combined to obtain richer representations of pedestrian regions, and better detection performance. We perform extensive evaluations on five different datasets to compare the performance of the proposed approach with two baselines. Moreover, we evaluate and compare the transferability of these approaches by doing leave-one-out cross validation across different datasets. Furthermore, the results show that the proposed approach outperforms the baselines, and has better transferability properties across different thermal image datasets.

47 OTHER INSTRUMENTATION↗

A susceptibility gene signature for ERBB2-driven mammary tumour development and metastasis in collaborative cross mice

Background: Deeper insights into ERBB2-driven cancers are essential to develop new treatment approaches for ERBB2+ breast cancers (BCs). We employed the Collaborative Cross (CC) mouse model to unearth genetic factors underpinning Erbb2-driven mammary tumour development and metastasis. Methods: 732 F1 hybrid female mice between FVB/N MMTV-Erbb2 and 30 CC strains were monitored for mammary tumour phenotypes. GWAS pinpointed SNPs that influence various tumour phenotypes. Multivariate analyses and models were used to construct the polygenic score and to develop a mouse tumour susceptibility gene signature (mTSGS), where the corresponding human ortholog was identified and designated as hTSGS. The importance and clinical value of hTSGS in human BC was evaluated using public datasets, encompassing TCGA, METABRIC, GSE96058, and I-SPY2 cohorts. The predictive power of mTSGS for response to chemotherapy was validated in vivo using genetically diverse MMTV-Erbb2 mice. Findings: Distinct variances in tumour onset, multiplicity, and metastatic patterns were observed in F1-hybrid female mice between FVB/N MMTV-Erbb2 and 30 CC strains. Besides lung metastasis, liver and kidney metastases emerged in specific CC strains. GWAS identified specific SNPs significantly associated with tumour onset, multiplicity, lung metastasis, and liver metastasis. Multivariate analyses flagged SNPs in 20 genes (Stx6, Ramp1, Traf3ip1, Nckap5, Pfkfb2, Trmt1l, Rprd1b, Rer1, Sepsecs, Rhobtb1, Tsen15, Abcc3, Arid5b, Tnr, Dock2, Tti1, Fam81a, Oxr1, Plxna2, and Tbc1d31) independently tied to various tumour characteristics, designated as a mTSGS. hTSGS scores (hTSGSS) based on their transcriptional level showed prognostic values, superseding clinical factors and PAM50 subtype across multiple human BC cohorts, and predicted pathological complete response independent of and superior to MammaPrint score in I-SPY2 study. The power of mTSGS score for predicting chemotherapy response was further validated in an in vivo mouse MMTV-Erbb2 model, showing that, like findings in human patients, mouse tumours with low mTSGS scores were most likely to respond to treatment. Interpretation: Our investigation has unveiled many new genes predisposing individuals to ERBB2-driven cancer. Translational findings indicate that hTSGS holds promise as a biomarker for refining treatment strategies for patients with BC.

60 APPLIED LIFE SCIENCES↗

Data analytics for leak detection in a subcritical boiler

For decades, boiler leaks have been the leading cause of forced outages in the coal-fired unit. The leak occurrences are currently escalating since the existing plants must satisfy faster-ramping rates to support grid operation. Data analytics including Principal Component Analysis, Canonical Variate, and Fisher Discriminant Analysis were combined for detecting and characterizing the leak in a commercial 650 MW subcritical coal-fired power plant. The combined approach was shown to be highly effective in the fault investigation that would not have been easily achieved by an individual technique. The variability in both training and validation datasets was first evaluated using PCA. Then, the CV-FDA was employed to discriminate among faults, and to categorize the processed data into two main groups: no-leak (0) and leak (1), providing the timeframe and location of the leak occurrence. Furthermore, about 8,014 observations from 81 process variables were initially included in the calculation, while the variable count was reduced to 4 with less than 1% misclassification rate in total observations. Finally, the leak was isolated in the waterwall section. Thus, the outcome of this research may provide early detection and isolation of faulty operations in the coal-fired power plant that involves a considerable number of process variables.

20 FOSSIL-FUELED POWER PLANTS↗

Tightly-coupled camera/LiDAR integration for point cloud generation from GNSS/INS-assisted UAV mapping systems

Unmanned aerial vehicles (UAVs) equipped with integrated global navigation satellite systems/inertial navigation systems (GNSS/INS) together with cameras and/or LiDAR sensors are being widely used for topographic mapping in a variety of applications such as precision agriculture, coastal monitoring, and archaeological documentation. Integration of image-based and LiDAR point clouds can provide a comprehensive 3D model of the area of interest. For such integration, ensuring a good alignment between data from the different sources is critical. Although many works have been conducted on this topic, there is still a need for a rigorous integration approach that minimizes the discrepancy between camera and LiDAR data caused by inaccurate system calibration parameters and/or trajectory artifacts. This study proposes an automated tightly-coupled camera/LiDAR integration workflow for GNSS/INS-assisted UAV systems. The proposed strategy is conducted in three main steps. First, an image-based point cloud is generated using a LiDAR/GNSS/INS-assisted structure from motion (SfM) strategy. Then, feature correspondences between image-based and LiDAR point clouds are automatically identified. Finally, an integrated-bundle adjustment procedure including image points, LiDAR raw measurements, and GNSS/INS information is conducted to minimize the discrepancy between point clouds from different sensors while estimating system calibration parameters and refining the trajectory information. The proposed SfM strategy and integration framework are evaluated using five datasets. The SfM results show that using LiDAR data can facilitate feature matching and further increase the number of reconstructed 3D points. The experimental results also illustrate that the developed automated camera/LiDAR integration strategy is capable of accurately estimating system calibration parameters to achieve good alignment among camera/LiDAR data from single/multiple systems. Finally, an absolute accuracy in the range of 3–5 cm is achieved for the image/LiDAR point clouds after the integration process.

42 ENGINEERING↗

Inverse mapping of properties to composition through generative modeling for designing molten salts

Generative modeling (GM) has been increasingly used for the inverse design and optimization of materials, yet its application to molten salt mixtures remains unexplored despite how a successful approach to the inverse design of molten salts would contribute to efficiently exploiting their customizability and unlocking their advantages in applications, such as energy production and energy storage. This work presents a workflow for the inverse design of molten salts with targeted density values, addressing the challenge of representing these complex mixtures in GM. A dataset of critically evaluated molten salt densities is used to train a variational autoencoder coupled with a predictive deep neural network, which then can be used to generate new molten salt compositions with desired density values. The effectiveness of the approach is demonstrated by designing mixtures with distinct densities and validating the predicted values using ab initio molecular dynamics simulations.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training

With the end of Moore’s law and Dennard scaling, efficient training increasingly requires rethinking data volume. Can we train better models with significantly less data via intelligent subsampling? To explore this, we develop SICKLE, a sparse intelligent curation framework for efficient learning, featuring a novel maximum entropy (MaxEnt) sampling approach, scalable training, and energy benchmarking. We compare MaxEnt with random and phase-space sampling on large direct numerical simulation (DNS) datasets of turbulence. Evaluating SICKLE at scale on Frontier, we show that subsampling as a preprocessing step can, in many cases, improve model accuracy and substantially lower energy consumption, with observed reductions of up to 38×.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

CYPminer: an automated cytochrome P450 identification, classification, and data analysis tool for genome data sets across kingdoms

Background: Cytochrome P450 monooxygenases (termed CYPs or P450s) are hemoproteins ubiquitously found across all kingdoms, playing a central role in intracellular metabolism, especially in metabolism of drugs and xenobiotics. The explosive growth of genome sequencing brings a new set of challenges and issues for researchers, such as a systematic investigation of CYPs across all kingdoms in terms of identification, classification, and pan-CYPome analyses. Such investigation requires an automated tool that can handle an enormous amount of sequencing data in a timely manner. Results: CYPminer was developed in the Python language to facilitate rapid, comprehensive analysis of CYPs from genomes of all kingdoms. CYPminer consists of two procedures i) to generate the Genome-CYP Matrix (GCM) that lists all occurrences of CYPs across the genomes, and ii) to perform analyses and visualization of the GCM, including pan-CYPomes (pan- and core-CYPome), CYP co-occurrence networks, CYP clouds, and genome clustering data. The performance of CYPminer was evaluated with three datasets from fungal and bacterial genome sequences. Conclusions: CYPminer completes CYP analyses for large-scale genomes from all kingdoms, which allows systematic genome annotation and comparative insights for CYPs. CYPminer also can be extended and adapted easily for broader usage.

59 BASIC BIOLOGICAL SCIENCES↗

2D reactive transport model of shale chemical weathering and biogeochemical fluxes along a mountainous hillslope, East River Watershed, Colorado: Input files and simulation results

This data package contains input files and simulation results for a two-dimensional (2D) reactive transport model used to quantitatively analyze the coupled hydrological and biogeochemical processes governing shale weathering and associated biogeochemical fluxes under realistic environmental conditions in the high-elevation East River Watershed. These data support the conclusions presented in Stolze et al. (Water Resources Research, under review), "Model-based interpretation of solute exports and carbon partitioning during shale weathering in a mountainous hillslope". The model simulates atmospheric-subsurface gas exchange, subsurface water flow, and shale weathering processes under dynamic, year-scale conditions along a shale-underlain hillslope located in the East River watershed. The simulations were performed using the PFLOTRAN flow and reactive transport code and executed on the Perlmutter supercomputer to leverage its large-scale parallel computing capabilities. The data package contains two zipped folders, "model_input_files" and "simulation_results", and one readme.txt file. "model_input_files" contains the necessary input files to run the calibrated base-base model presented in Stolze et al. (Water Resources Research, under review). "simulation_results" contains a single hdf5 file ("Output_2D_hillslope_model.h5") which includes the results of simulation performed using the base-case model. This file can be opened with HDFView 3.1.4, Python, or MATLAB. "readme.txt" contains relevant information about the base-case model and provides guidelines on how to run the associated input files provided in the folder "model_input_files". Furthermore, readme.txt provides information regarding the model results provided in "Output_2D_hillslope_model.h5" such as matrix dimensionality and output units. Field datasets used to evaluate model performance were collected at three monitoring wells located along a hillslope transect (PLM1, PLM2, and PLM3). Dissolved ion concentration data were collected from November 2016 to October 2021 for Ca, Mg, DIC, Na, K, SO4 (Dong et al., 2025 - dic_npoc_data_2014_2024.zip - DOI:10.15485/1660459; Williams et al., 2025 - anion_data_2014_2024.zip - DOI:10.15485/1668054; Dong et al., 2025 - cation_data_2014_2024.zip - DOI:10.15485/1668055). Note that we used the files named er_PLM1_xx_yy, er_PLM2_xx_yy, and er_PLM3_xx_yy where xx stands for the name of the aqueous species and yy stands for the depth where the measurements were performed. Soil water content ([0 - 1] m) and water table depth were collected from November 2016 to October 2021 (Wan et al., 2024 - Dynamic_water_table__depthsFig2b.csv and Soil_water_content_Fig4e.csv - DOI:10.15485/2322567). Gaseous CO2 concentration were collected from October 2020 to December 2021(Wan et al., 2024 - Soil_CO2_concentrations_Fig4h.csv - DOI:10.15485/2322567) Gaseous CO2 flux from the subsurface to the atmosphere were collected in the vicinity of PLM2 from October 2019 to May 2022 (Wu et al., 2025). Soil microbial biomass concentration was measured from August 2016 to June 2017 (Sorensen et al., 2019 - 2017_East_River_Pumphouse_Microbial_Biomass__1_.csv - DOI:10.15485/1577267) All field data are published as CSV files compatible with Microsoft Excel, MATLAB, and Python, or as text files. The coordinates of the monitoring wells and the CO2(g) flux sensor in the coordinate system WGS84 are: -PLM1: [38.9197710 ; -106.9492750] -PLM2: [38.9201580 ; -106.9487170] -PLM3: [38.9207843 ; -106.9483668] -PLM4: 38.9210060 ; -106.9479528] -CO2(g) flux sensor: [38.9199180 ; -106.9489906] ------------------------------------------------------------------------------------------- This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research used resources of the National Energy Research Scientific Computing Center (NERSC), a Department of Energy User Facility using NERSC award BER-ERCAP 23980, BER-ERCAP 28550, and BER-ERCAP 33789.

54 ENVIRONMENTAL SCIENCES↗

Gradient-Based Novelty Detection Boosted by Self-Supervised Binary Classification

Novelty detection aims to automatically identify out-of-distribution (OOD) data, without any prior knowledge of them. It is a critical step in data monitoring, behavior analysis and other applications, helping enable continual learning in the field. Conventional methods of OOD detection perform multi-variate analysis on an ensemble of data or features, and usually resort to the supervision with OOD data to improve the accuracy. In reality, such supervision is impractical as one cannot anticipate the anomalous data. In this paper, we propose a novel, self-supervised approach that does not rely on any pre-defined OOD data: (1) The new method evaluates the Mahalanobis distance of the gradients between the in-distribution and OOD data. (2) It is assisted by a self-supervised binary classifier to guide the label selection to generate the gradients, and maximize the Mahalanobis distance. In the evaluation with multiple datasets, such as CIFAR-10, CIFAR-100, SVHN and TinyImageNet, the proposed approach consistently outperforms state-of-the-art supervised and unsupervised methods in the area under the receiver operating characteristic (AUROC) and area under the precision-recall curve (AUPR) metrics. We further demonstrate that this detector is able to accurately learn one OOD class in continual learning.

Sun, Jingbo↗

An Exact Algorithm for the Linear Tape Scheduling Problem

Magnetic tapes are often considered as an outdated storage technology, yet they are still used to store huge amounts of data. Their main interests are a large capacity and a low price per gigabyte, which come at the cost of a much larger file access time than on disks. With tapes, finding the right ordering of multiple file accesses is thus key to performance. Moving the reading head back and forth along a kilometer long tape has a non-negligible cost and unnecessary movements thus have to be avoided. However, the optimization of tape request ordering has rarely been studied in the scheduling literature, much less than I/O scheduling on disks. For instance, minimizing the average service time for several read requests on a linear tape remains an open question. Therefore, in this paper, we aim at improving the quality of service experienced by users of tape storage systems, and not only the peak performance of such systems. To this end, we propose a reasonable polynomial-time exact algorithm while this problem and simpler variants have been conjectured NP-hard. We also refine the proposed model by considering U-turn penalty costs accounting for inherent mechanical accelerations. Then, we propose a low-cost variant of our optimal algorithm by restricting the solution space, yet still yielding an accurate suboptimal solution. Finally, we compare our algorithms to existing solutions from the literature on logs of the mass storage management system of a major datacenter. This allows us to assess the quality of previous solutions and the improvement achieved by our low-cost algorithm. Aiming for reproducibility, we make available the complete implementation of the algorithms used in our evaluation, alongside the dataset of tape requests that is, to the best of our knowledge, the first of its kind to be publicly released.

Honoré, Valentin↗

National Virtual Biotechnology Laboratory: Report on Rapid R&D Solutions to the COVID-19 Crisis

With funding from the CARES Act, the U.S Department of Energy (DOE) established the National Virtual Biotechnology Laboratory (NVBL) in March 2020 to address key challenges associated with the COVID-19 crisis. NVBL brought together the broad scientific and technical expertise and resources of DOE’s 17 national laboratories to help tackle medical supply short ages, discover potential drugs to fight the virus, develop and validate COVID-19 testing methods, model disease spread and impact across the nation, and understand virus transport in buildings and the environment. National laboratory resources leveraged for this effort include a suite of world-leading user facilities broadly available to the research community, such as light and neutron sources, nanoscale science research centers, sequencing and biocharacterization facilities, and high-performance computing facilities. Within months, NVBL teams produced innovations in materials and advanced manufacturing that mitigated shortages in test kits and personal protective equipment (PPE), creating nearly 1,000 new jobs. They used DOE’s high-performance computers and light and neutron sources to identify promising candidates for antibodies and antivirals that universities and drug companies are now evaluating. NVBL researchers also developed new diagnostic targets and sample collection approaches, and supported U.S. Food and Drug Administration (FDA), Centers for Disease Control and Prevention (CDC), and U.S. Department of Defense (DoD) efforts to establish national guidelines used in administering millions of tests. Researchers used artificial intelligence and high-performance computing to produce near-real-time data analysis to forecast disease transmission, stress on public health infrastructure, and economic impact, which supported decision-makers at the local, state, and national levels. NVBL teams also studied how to control indoor virus movement to minimize uptake and protect human health. NVBL’s accomplishments demonstrate not only the powerful resource represented by DOE’s national laboratories working together to meet national needs, but also the effectiveness of the integrated NVBL framework for rapidly responding to emergencies with research and development (R&D) solutions. As the fight against COVID continues, sustained efforts are needed to confront this pandemic as well as future threats. Examples include: 1) Establishing “supply chains on demand” to meet emergency production needs by leveraging the materials and manufacturing expertise of DOE national laboratories and developing advances in electronics, sensing, robotics, and automation capabilities; 2) Improving the speed and robustness of drug discovery by integrating experimental platforms with DOE’s computational and experimental user facilities, which provide unique resources to support the discovery of high-potential therapeutic agents; 3) Protecting public, environmental, and animal health by developing new testing protocols and instrumentation adaptable to diverse sample types (both physiological and environmental) to quickly detect a wide range of pathogens and monitor other biorisks; 4) Supporting near-real-time data needs of decision-makers at the local, regional, state, and national levels by advancing data curation, analysis, and modeling using artificial intelligence and new data science tools for managing and evaluating large diverse datasets; 5) Harnessing DOE’s expertise in environmental modeling to design rooms and air handling for offices, classrooms, restaurants, and other structures to minimize biorisk transmissions. Going forward, NVBL is poised to apply the unique capabilities and expertise of the national laboratory complex to future national and international emergencies, both natural and engineered. Through this framework, the Office of Science will continue to be an integral component of agency wide efforts to prepare for and respond to biorisks and other crises.

42 ENGINEERING↗

Rapid identification of enteric bacteria from whole genome sequences using average nucleotide identity metrics

Identification of enteric bacteria species by whole genome sequence (WGS) analysis requires a rapid and an easily standardized approach. We leveraged the principles of average nucleotide identity using MUMmer (ANIm) software, which calculates the percent bases aligned between two bacterial genomes and their corresponding ANI values, to set threshold values for determining species consistent with the conventional identification methods of known species. The performance of species identification was evaluated using two datasets: the Reference Genome Dataset v2 (RGDv2), consisting of 43 enteric genome assemblies representing 32 species, and the Test Genome Dataset (TGDv1), comprising 454 genome assemblies which is designed to represent all species needed to query for identification, as well as rare and closely related species. The RGDv2 contains six Campylobacter spp., three Escherichia/Shigella spp., one Grimontia hollisae, six Listeria spp., one Photobacterium damselae, two Salmonella spp., and thirteen Vibrio spp., while the TGDv1 contains 454 enteric bacterial genomes representing 42 different species. The analysis showed that, when a standard minimum of 70% genome bases alignment existed, the ANI threshold values determined for these species were ≥95 for Escherichia/Shigella and Vibrio species, ≥93% for Salmonella species, and ≥92% for Campylobacter and Listeria species. Using these metrics, the RGDv2 accurately classified all validation strains in TGDv1 at the species level, which is consistent with the classification based on previous gold standard methods.

59 BASIC BIOLOGICAL SCIENCES↗