Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Reference call set”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Speeding genomic island discovery through systematic design of reference database composition

Background Genomic islands (GIs) are mobile genetic elements that integrate site-specifically into bacterial chromosomes, bearing genes that affect phenotypes such as pathogenicity and metabolism. GIs typically occur sporadically among related bacterial strains, enabling comparative genomic approaches to GI identification. For a candidate GI in a query genome, the number of reference genomes with a precise deletion of the GI serves as a support value for the GI. Our comparative software for GI identification was slowed by our original use of large reference genome databases (DBs). Here we explore smaller species-focused DBs. Results With increasing DB size, recovery of our reliable prophage GI calls reached a plateau, while recovery of less reliable GI calls (FPs) increased rapidly as DB sizes exceeded ~500 genomes; i.e., overlarge DBs can increase FP rates. Paradoxically, relative to prophages, FPs were both more frequently supported only by genomes outside the species and more frequently supported only by genomes inside the species; this may be due to their generally lower support values. Setting a DB size limit for our SMA ll R anked T ailored (SMART) DB design speeded runtime ~65-fold. Strictly intra-species DBs would tend to lower yields of prophages for small species (with few genomes available); simulations with large species showed that this could be partially overcome by reaching outside the species to closely related taxa, without an FP burden. Employing such taxonomic outreach in DB design generated redundancy in the DB set; as few as 2984 DBs were needed to cover all 47894 prokaryotic species. Conclusions Runtime decreased dramatically with SMART DB design, with only minor losses of prophages. We also describe potential utility in other comparative genomics projects.

59 BASIC BIOLOGICAL SCIENCES↗

CORPSE model with litter decomposition parameters derived from the LIDET dataset

This is a version of the CORPSE model (Carbon, Organisms, Rhizosphere and Protection in the Soil Environment, Sulman et al. 2014) that uses litter decomposition parameters derived from a modified Monte Carlo simulation using the LIDET litter decomposition dataset (Long-term Intersite Decomposition Experiment Team, Harmon 2013). The code also includes the Baseline parameters, and the eight other best parameter sets identified in a modified Monte Carlo simulation. Related publication:Juice, S.M., Ridgeway, J.R., Hartman, M.D., Parton, W.J., Berardi, D.M., Sulman, B.N., Allen, K.E., & Brzostek, E.R. Reparameterizing litter decomposition using a simplified Monte Carlo method improves litter decay simulated by a microbial model and alters bioenergy soil carbon estimates. Description of files:The folder "Input Files" contains one folder for each LIDET site with data necessary to run the model. Note that "(site)" in the filenames below indicates where the LIDET site code appears (see Table 1 for site codes). Data streams include: CORPSE_full_spinup_litter.csv, CORPSE_full_spinup_rhizo.csv, CORPSE_full_spinup_bulk.csv, litterbag_init_100g_6spp.csv: initial C and N (kg C or N/m2) pool values for each soil layer, the litterbag_init_100_6spp.csv file is for the litterbag layer and is the same file for all sites. All initial C and N files have the same columns (Column - Description - Units) uFastC - Unprotected fast decomposing carbon - kg carbon/m2 uSlowC - Unprotected slow decomposing carbon - kg carbon/m2 uNecroC - Unprotected necromass carbon - kg carbon/m2 pFastC - Protected fast decomposing carbon - kg carbon/m2 pSlowC - Protected slow decomposing carbon - kg carbon/m2 pNecroC - Protected necromass carbon - kg carbon/m2 livingMicrobeC - Carbon in living microbial biomass - kg carbon/m2 uFastN - Unprotected fast decomposing nitrogen - kg nitrogen/m2 uSlowN - Unprotected slow decomposing nitrogen - kg nitrogen/m2 uNecroN - Unprotected necromass nitrogen - kg nitrogen/m2 pFastN - Protected fast decomposing nitrogen - kg nitrogen/m2 pSlowN - Protected slow decomposing nitrogen - kg nitrogen/m2 pNecroN - Protected necromass nitrogen - kg nitrogen/m2 inorganicN - Inorganic nitrogen - kg nitrogen/m2 CO2 - Carbon in carbon dioxide - kg carbon/m2 livingMicrobeN - Nitrogen in living microbial biomass - kg nitrogen/m2 soilT (site) DOY274start.csv: Average daily soil temperature (oC) interpolated from previously calculated monthly values used in DayCent LIDET simulations (Bonan et al., 2013). soilT (site) DOY274start.csv: Average daily soil volumetric water content (VWC) scalar interpolated from previously calculated monthly values used in DayCent LIDET simulations (Bonan et al., 2013). litter production.csv: Average daily litter production values for each site, data sources listed in Table S3 of related publication. litter (site) CN.csv: C:N ratio for each species from LIDET dataset (Table 2, Harmon 2013). (site).csv: Table indicating number of observations for each species decomposed at each site. Instructions: Save the model code ("CORPSE_LIDET.R") and "Input Files" folder in the same folder. Also make a folder for the model output (e.g., "results_Baseline") in the same folder. Set the working directory (setwd) in the model code to the folder with the files saved in step #1. Select the parameter set to use for the litter and litterbag compartments, comment out all other parameter sets. Run code. Output will be saved in the folder made in step 1. Output destination can be changed as necessary in code section called "Running the model." Table 1 LIDET sites and site codes used in model files. Site Code - Site AND - H.J. Andrews Experimental Forest BNZ - Bonanza Creek Experimental Forest BSF - Blodgett Research Forest CDR - Cedar Creek Natural History Area CPR - Central Plains Experimental Range HBR - Hubbard Brook Experimental Forest HFR - Harvard Forest JUN - Juneau KBS - Kellogg Biological Station KNZ - Konza Prairie Research Natural Area NWT - Niwot Ridge/Green Lakes Valley OLY - Olympic National Park OLY Conifer forest SEV - Sevilleta National Wildlife Refuge SMR - Santa Margarita Ecological Reserve UFL - University of Florida VCR - Virginia Coast Reserve Table 2 LIDET species and species codes used in model files (6 common species). Species - Species Code Sugar maple (Acer saccharum) - ACSA Drypetes (Drypetes glauca) - DRGL Red pine (Pinus resinosa) - PIRE Chestnut oak (Quercus prinus) - QUPR Western redcedar (Thuja plicata) - THPL Wheat (Triticum aestivum) - TRAE References:Bonan, G. B., Hartman, M. D., Parton, W. J., & Wieder, W. R. (2013). Evaluating litter decomposition in earth system models with long-term litterbag experiments: an example using the Community Land Model version 4 (CLM4). Global Change Biology, 19(3), 957-974. https://doi.org/https://doi.org/10.1111/gcb.12031 Harmon, M. (2013). LTER Intersite Fine Litter Decomposition Experiment (LIDET), 1990 to 2002. Long-Term Ecological Research. Forest Science Data Bank, Corvallis, OR. [Data set]. Accessed http://andlter.forestry.oregonstate.edu/data/abstract.aspx?dbcode=TD023. https://doi.org/10.6073/pasta/f35f56bea52d78b6a1ecf1952b4889c5. Sulman, B. N., Phillips, R. P., Oishi, A. C., Shevliakova, E., & Pacala, S. W. (2014). Microbe-driven turnover offsets mineral-mediated storage of soil carbon under elevated CO2. Nature Climate Change, 4, 1099 - 1102. https://doi.org/10.1038/nclimate2436

Juice, Stephanie↗

Measuring Chemical Likeness of Stars with Relevant Scaled Component Analysis

Identification of chemically similar stars using elemental abundances is core to many pursuits within Galactic archeology. However, measuring the chemical likeness of stars using abundances directly is limited by systematic imprints of imperfect synthetic spectra in abundance derivation. We present a novel data-driven model that is capable of identifying chemically similar stars from spectra alone. We call this relevant scaled component analysis (RSCA). RSCA finds a mapping from stellar spectra to a representation that optimizes recovery of known open clusters. By design, RSCA amplifies factors of chemical abundance variation and minimizes those of nonchemical parameters, such as instrument systematics. The resultant representation of stellar spectra can therefore be used for precise measurements of chemical similarity between stars. We validate RSCA using 185 cluster stars in 22 open clusters in the Apache Point Observatory Galactic Evolution Experiment survey. We quantify our performance in measuring chemical similarity using a reference set of 151,145 field stars. We find that our representation identifies known stellar siblings more effectively than stellar-abundance measurements. Using RSCA, 1.8% of pairs of field stars are as similar as birth siblings, compared to 2.3% when using stellar-abundance labels. We find that almost all of the information within spectra leveraged by RSCA fits into a two-dimensional basis, which we link to [Fe/H] and α-element abundances. We conclude that chemical tagging of stars to their birth clusters remains prohibitive. However, using the spectra has noticeable gain, and our approach is poised to benefit from larger data sets and improved algorithm designs.

79 ASTRONOMY AND ASTROPHYSICS↗

RAG for FLAG: AI Assistance for a Physics Code

Artificial intelligence (AI) has quickly become an important tool in scientific research, where significant efforts are underway to develop tools that will expedite the research process. One area of particular impact is scientific software, which can be particularly complex, and therefore time consuming to learn and use effectively. AI assistants are increasingly helping to streamline the process by performing tasks such as interactively answering user questions or suggesting solutions. Los Alamos National Laboratory (LANL) develops several advanced scientific codes, such as FLAG, which can be used to run multiphysics simulations. With this study, our goal was to develop an AI assistant for FLAG that could help make the process of understanding the software and running physics simulations more efficient. To develop an AI assistant for FLAG, we used a method called retrieval-augmented generation (RAG), which is a technique that uses information from relevant data sources to enhance the accuracy of large language models (LLMs). We used the FLAG user manual and other FLAG documentation as the knowledge base for the RAG system. When a user provides a query, RAG retrieves relevant sections from the knowledge base in response, then uses those excerpts to generate grounded and contextually rich answers. We found that our AI assistant was able to provide context aware answers and source references to user queries. To evaluate performance, we developed a set of 40 benchmark questions and compared the accuracy of the responses to those of two standard LLMs without retrieval. Our AI assistant significantly outperformed the standard LLMs at answering FLAG-related questions, with an 82.5% accuracy rate, compared to 47.5% for both of the standard LLMs. This has the potential to make the process of learning and using FLAG much easier, especially for new users. Ultimately, it supports LANL’s broader mission by empowering scientists and engineers to focus more on discovery and analysis rather than on navigating complex software systems.

97 MATHEMATICS AND COMPUTING↗

Geothermal Play Fairway Analysis, Part 2: GIS methodology

Play Fairway Analysis (PFA) in geothermal exploration originates from a systematic methodology developed within the petroleum industry and is based on a geologic, geophysical, and hydrologic framework of identified geothermal systems. We tailored this methodology to study the geothermal resource potential of the Snake River Plain and surrounding region, but it can be adapted to other geothermal resource settings. We adapted the PFA approach to geothermal resource exploration by cataloging the critical elements controlling exploitable hydrothermal systems, establishing risk matrices that evaluate these elements in terms of both probability of success and level of knowledge, and building a code-based ‘processing model’ to process results. A geographic information system was used to compile a range of different data types, which we refer to as elements (e.g., faults, vents, heat flow, etc.), with distinct characteristics and measures of confidence. Discontinuous discrete data (points, lines, or polygons) for each element were transformed into continuous interpretive 2D grid surfaces called evidence layers. Because different data types have varying uncertainties, most evidence layers have an accompanying confidence layer which reflects spatial variations in these uncertainties. Risk layers, as defined here, are the product of evidence and confidence layers, and are the building blocks used to construct Common Risk Segment (CRS) maps for heat, permeability, and seal, using a weighted sum for permeability and heat, but a different approach with seal. CRS maps quantify the variable risk associated with each of these critical components. In a final step, the three CRS maps were combined into a Composite Common Risk Segment (CCRS) map, using a modified weighted sum, for results that reveal favorable areas for geothermal exploration. Additional maps are also presented that do not mix contributions from evidence and confidence (to allow an isolated view of evidence and confidence), as well as maps that calculate favorability using the product of components instead of a weighted sum (to highlight where all components are present). Our approach helped to identify areas of high geothermal favorability in the western and central Snake River Plain during the first phase of study and helped identify more precise local drilling targets during the second phase of work. By identifying favorable areas, this methodology can help to reduce uncertainty in geothermal energy exploration and development.

15 GEOTHERMAL ENERGY↗

Advanced Energy Scale Correction Techniques for the X-ray Transition Edge Sensors of the Athena mission

The X-ray Integral Field Unit (X-IFU) onboard the future European X-ray telescope Athena will be the first space instrument carrying an array of more than a thousand transition edge sensors. One of the key challenges of the X-IFU is the measurement of narrow X-ray atomic lines to determine velocity shifts at an unprecedented level of accuracy. For this reason, the energy scale of the instrument needs to be known with extreme accuracy, of 0.4 eV (1σ) up to 7 keV. The energy scale will be measured on the ground through a dedicated calibration campaign using fiducial X-ray sources. Though calibrated, the energy scale is extremely sensitive to the environmental conditions around the TES array, and drifts in the readout chain electronics. Uncorrected, the energy scale can naturally drift up to hundreds of eVs. Changes of the TES gain will be monitored via onboard X-ray calibration sources, and the energy scale will be corrected either per pixel, or within a small groups of pixels. Although simulations show that a 0.4 eV level can be achieved, the very high accuracy required by the X-IFU calls for experimental validation. A dedicated measurement campaign has been performed by NASA Goddard Space Flight Center to characterize the energy scale of a prototype kilo-pixel array of X-IFU-representative TESs. The analysis of the data demonstrated the ability to correct for various drifts using two fiducial lines to track the temporal gain variation. In this paper, we propose to extend this study on the same data set by investigating multi-parameter correction techniques based on both the pulse-height of the fiducial line and the prepulse baseline level, using the knowledge of the TES energy scale at reference temperature/magnetic field set points acquired on the ground. Investigations on the co-adding of pixels to perform a joint correction over pools of pixels is also explored.

79 ASTRONOMY AND ASTROPHYSICS↗

Solution of the Schrödinger equation for quasi-one-dimensional materials using helical waves

We formulate and implement a spectral method for solving the Schrödinger equation, as it applies to quasi-one-dimensional materials and structures. This allows for computation of the electronic structure of important technological materials such as nanotubes (of arbitrary chirality), nanowires, nanoribbons, chiral nanoassemblies, nanosprings and nanocoils, in an accurate, efficient and systematic manner. Our work is motivated by the observation that one of the most successful methods for carrying out electronic structure calculations of bulk/crystalline systems — the plane-wave method — is a spectral method based on eigenfunction expansion. Our scheme avoids computationally onerous approximations involving periodic supercells often employed in conventional plane-wave calculations of quasi-one-dimensional materials, and also overcomes several limitations of other discretization strategies, e.g., those based on finite differences and atomic orbitals. The basis functions in our method — called helical waves (or twisted waves) — are eigenfunctions of the Laplacian with symmetry adapted boundary conditions, and are expressible in terms of plane waves and Bessel functions in helical coordinates. We describe the setup of fast transforms to carry out discretization of the governing equations using our basis set, and the use of matrix-free iterative diagonalization to obtain the electronic eigenstates. Miscellaneous computational details, including the choice of eigensolvers, use of a preconditioning scheme, evaluation of oscillatory radial integrals and the imposition of a kinetic energy cutoff are discussed. We have implemented these strategies into a computational package called HelicES (Helical Electronic Structure). We demonstrate the utility of our method in carrying out systematic electronic structure calculations of various quasi-one-dimensional materials through numerous examples involving nanotubes, nanoribbons and nanowires. We also explore the convergence properties of our method, and assess its accuracy and computational efficiency by comparison against reference finite difference, transfer matrix method and plane-wave results. We anticipate that our method will find applications in computational nanomechanics and multiscale modeling, for carrying out transport calculations of interest to the field of semiconductor devices, and for the discovery of novel chiral phases of matter that are of relevance to the burgeoning quantum hardware industry.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Water Level Data from Wells PLM1 and PLM6 for the East River Watershed, Colorado

This dataset (Williams et al., 2020) contains the original un-QA/QC-ed water level data for PLM1 and PLM6 and has been obsoleted. The data contained within this dataset is not to be used. Refer to Faybishenko et al., 2022 (DOI: 10.15485/1866836) for the latest QA/QC-ed data available via ESS-DIVE.This data set contains water level data for the PLM1 and PLM6 wells. PLM1 and PLM6 are location identifiers used by the Watershed Function SFA project for two groundwater monitoring wells along an elevation gradient located along the lower montane life zone of a hillslope near the Pumphouse location. These wells used to monitor subsurface water and carbon inventories and fluxes at the East River Watershed, Colorado, USA. Complete metadata information on the PLM1 and PLM6 wells are available in the related data package reference Varadharajan C, et al (2020). https://doi.org/10.15485/1660962.Data are reported in .csv files per well. The latitude and longitude of each location are given in a file called locations.csv. These data are used for determining the seasonally dependent flow of groundwater under the PLM hillslope. The downslope flow of groundwater in combination with data on groundwater chemistry can be used to estimate rates of solute export from the hillslope to the floodplain and river.These data products are part of the Watershed Function Scientific Focus Area collection effort to further scientific understanding of biogeochemical dynamics from genome to watershed scales.

54 ENVIRONMENTAL SCIENCES↗

Subtleties in the trainability of quantum machine learning models

A new paradigm for data science has emerged, with quantum data, quantum models, and quantum computational devices. This field, called quantum machine learning (QML), aims to achieve a speedup over traditional machine learning for data analysis. However, its success usually hinges on efficiently training the parameters in quantum neural networks, and the field of QML is still lacking theoretical scaling results for their trainability. Some trainability results have been proven for a closely related field called variational quantum algorithms (VQAs). While both fields involve training a parametrized quantum circuit, there are crucial differences that make the results for one setting not readily applicable to the other. In this work, we bridge the two frameworks and show that gradient scaling results for VQAs can also be applied to study the gradient scaling of QML models. Our results indicate that features deemed detrimental for VQA trainability can also lead to issues such as barren plateaus in QML. Consequently, our work has implications for several QML proposals in the literature. In addition, we provide theoretical and numerical evidence that QML models exhibit further trainability issues not present in VQAs, arising from the use of a training dataset. We refer to these as dataset-induced barren plateaus. These results are most relevant when dealing with classical data, as here the choice of embedding scheme (i.e., the map between classical data and quantum states) can greatly affect the gradient scaling.

97 MATHEMATICS AND COMPUTING↗

Out of Distribution Detection with Neural Network Anchoring

This is code to reproduce and build on OOD detection from the paper "Out of Distribution Detection with Neural Network Anchoring". Our goal here is to exploit heteroscedastic temperature scaling as a calibration strategy for out of distribution (OOD) detection. Heteroscedasticity here refers to the fact that the optimal temperature parameter for each sample can be different, as opposed to conventional approaches that use the same value for the entire distribution. To enable this, we propose a new training strategy called anchoring that can estimate appropriate temperature values for each sample, leading to state-of-the-art OOD detection performance across several benchmarks. Using NTK theory, we show that this temperature function estimate is closely linked to the epistemic uncertainty of the classifier, which explains its behavior. In contrast to some of the best-performing OOD detection approaches, our method does not require exposure to additional outlier datasets, custom calibration objectives, or model ensembling. Through empirical studies with different OOD detection settings - far OOD, near OOD, and semantically coherent OOD - we establish a highly effective OOD detection approach.

Thiagarajan, Jayaraman↗

A discontinuous piecewise polynomial generalized moving least squares scheme for robust finite element analysis on arbitrary grids

A variational approach is developed with a meshless discretization to enable accurate and robust numerical simulation of partial differential equations for meshes that are of poor quality. Traditional finite element methods use the mesh to both discretize the geometric domain and to define the finite element shape functions. The latter creates a dependence between the quality of the mesh and the properties of the finite element basis that may adversely affect the accuracy of the discretized problem. Here, we propose a new approach for defining finite element shape functions that breaks this dependence and separates mesh quality from the discretization quality, which we call discontinuous piecewise polynomial generalized moving least squares (DPP-GMLS). At the core of the approach is a meshless definition of the shape functions, which limits the purpose of the mesh to representing the geometric domain and integrating the basis functions without having any role in their approximation quality. The resulting non-conforming space can be utilized within a standard discontinuous Galerkin framework, providing a rigorous foundation for solving partial differential equations on low-quality meshes. We present a collection of numerical experiments demonstrating our approach in a wide range of settings: strongly coercive elliptic problems, linear elasticity in the compressible regime, and the stationary Stokes problem. We demonstrate convergence for all problems and stability for element pairs for problems which usually require inf-sup compatibility for conforming methods, also referring to a minor modification possible through the symmetric interior penalty Galerkin framework for stabilizing element pairs that would otherwise be traditionally unstable. Mesh robustness is particularly critical for elasticity, and we provide an example that our approach provides a greater than 5 x improvement in accuracy and allows for taking an 8 x larger stable timestep for a highly deformed mesh, compared to the continuous Galerkin finite element method.

97 MATHEMATICS AND COMPUTING↗

DL-TODA: A Deep Learning Tool for Omics Data Analysis

Metagenomics is a technique for genome-wide profiling of microbiomes; this technique generates billions of DNA sequences called reads. Given the multiplication of metagenomic projects, computational tools are necessary to enable the efficient and accurate classification of metagenomic reads without needing to construct a reference database. The program DL-TODA presented here aims to classify metagenomic reads using a deep learning model trained on over 3000 bacterial species. A convolutional neural network architecture originally designed for computer vision was applied for the modeling of species-specific features. Using synthetic testing data simulated with 2454 genomes from 639 species, DL-TODA was shown to classify nearly 75% of the reads with high confidence. The classification accuracy of DL-TODA was over 0.98 at taxonomic ranks above the genus level, making it comparable with Kraken2 and Centrifuge, two state-of-the-art taxonomic classification tools. DL-TODA also achieved an accuracy of 0.97 at the species level, which is higher than 0.93 by Kraken2 and 0.85 by Centrifuge on the same test set. Application of DL-TODA to the human oral and cropland soil metagenomes further demonstrated its use in analyzing microbiomes from diverse environments. Compared to Centrifuge and Kraken2, DL-TODA predicted distinct relative abundance rankings and is less biased toward a single taxon.

59 BASIC BIOLOGICAL SCIENCES↗

NRAP-Open-IAM Analytical Reservoir Model: Development and Testing

Geological carbon sequestration (GCS) is a key technology for reducing global carbon dioxide (CO 2 ) emissions. Over the last decade, the U.S. Department of Energy has invested in understanding the science base, developing practical implementation methods, and demonstrating secure GCS technologies to mitigate the environmental impacts associated with the atmospheric release of CO 2 . As part of the National Risk Assessment Partnership, a systems-level risk assessment tool, called the NRAP-Open-IAM, has been developed to conduct risk assessment and enable safe operations at a GCS site. The current NRAP-Open-IAM contains a simple reservoir model component that calculates the evolution of CO 2 saturation and fluid pressure in a storage reservoir during CO 2 injection operations. This report presents the development and testing of a new analytical reservoir reduced-order model (ROM), which is extended from an existing semi-analytical model for estimation of CO 2 and brine leakage along legacy wells, and enhances the capability of the NRAP-Open-IAM to simulate more types of reservoir conditions. The developed model is validated against three reference studies, and the results indicate that the new ROM predicts the behavior of the two-phase fluids (brine and injected CO 2 ) well and is applicable to different reservoir simulation boundary conditions (i.e., constant pressure boundary and infinite-acting boundary) without a priori user specification of the boundary type. Sensitivity analysis for a set of model parameters is performed using 4,000 synthetic cases prepared via a fully automated process and using machine-learning-based feature selection. The stochastic analysis identifies gravitational number (i.e., ratio of gravitational forces to viscous force) and distance between the injection well and observation location as the most impactful parameters for matching the pressure and CO 2 saturation, respectively, between the numerical simulations and the ROM. This report details the possible ROM uncertainties and serves as a guide for users to understand the use and limitations of this ROM. The code implementation of the model will be released as a module within the NRAP-Open-IAM.

54 ENVIRONMENTAL SCIENCES↗

ORNL_AISD_NiNb

This dataset describes the nickel-niobium solid solution binary alloy, where the two constituent elements nickel (Ni) and niobium (Nb) are randomly placed on an underlying crystal lattice. This dataset for nickel-niobium (Ni-Nb) alloys available includes the formation energy and bulk modulus for each crystal structure. Each atomic sample has a disordered phase which is obtained starting from an initial regular crystal structure of type body-centered cubic (BCC), face-centered cubic (FCC), or hexagonal compact packed (HCP). The geometry optimization ensures that all the alloy samples reached the equilibrium with negative formation energy. We perform geometry optimizations using the LAMMPS simulation package [1], a flexible simulation tool for particle-based materials modeling at the atomic, meso, and continuum scales. We utilized the embedded atom model (EAM) potential for Ni and Nb developed in a previous study [2]. The potential could describe behaviors of the liquid and solid phases of Ni-Nb alloy. The structural factors and angular distributions of three atoms are well-matched with X-ray and ab initio-based molecular dynamics data. We prepared the three different crystals with different initial lattice parameters (3.52 Ã… for FCC, 3.32 Ã… for BCC, and 3.5 Ã… for HCP). We performed energy minimization in two steps. Firstly, we minimized the structures with an isotropic unit cell to minimize the side effects from our arbitrary lattice parameters for all other compositions. Then, we applied geometry optimization with a triclinic (non-orthogonal) unit cell to fully minimize the stress components to calculate the elastic constants. In this procedure, we chose 10,000 as the maximum number of allowable steps aimed at obtaining fully relaxed atomic geometries. The dataset consists of three sets of crystal structures. The first set contains 46,086 irregular crystal structures, each of them with 54 atoms, obtained through optimization starting from a regular BCC crystal structure. The second set contains 24,543 irregular crystal structures, each of them with 32 atoms, obtained through optimization starting from a regular FCC crystal structure. The third set contains 39,303 irregular crystal structures, each of them with 48 atoms, obtained through optimization starting from a regular HCP crystal structure. The atomic configurations within each set span the possible compositional range. The three sets have been unified in a global dataset, which is extremely heterogeneous in terms of crystal structures, lattice volumes, and atomic configurations. Organization of files inside the dataset: the dataset contains three subdirectories called • BCC_opt • FCC_opt • HCP_opt based on the type of initial regular structure used to start the geometry optimization. Inside each of these folders, every atomic structure is identified by a string “A_B_Câ€, where A denotes the number of Nb in the system, B denotes index of structure with a given Nb number, and C denotes the total number of structures generated with a given Nb number. For each optimized crystal structure identified by the unique string of characters “A_B_Câ€, three files are provided: • A_B_C_opt.xyz: The optimized geometries in xyz format • A_B_C_opt.cfg: The optimized geometries in cfg format. It includes cell information and atomic energy, and forces calculated from LAMMPS. • A_B_C.elastic: Raw data of 21 elastic constants from LAMMPS output. • A_B_C.bulk: Calculated upper and lower bounds of bulk modulus and averaged one based on Voigt-Reuss-Hill approach from *.elastic. References: [1] A. P. Thompson, H. M. Aktulga, R. Berger, D. S. Bolintineanu, W. M. Brown, P. S. Crozier, P. J. in 't Veld, A. Kohlmeyer, S. G. Moore, T. D. Nguyen, R. Shan, M. J. Stevens, J. Tranchida, C. Trott, and S. J. Plimpton. LAMMPS - a flexible simulation tool for particle-based materials modeling at the atomic, meso, and continuum scales. Comp. Phys. Comm., 271:108171, 2022. [2] Y Zhang, R Ashcraft, MI Mendelev, CZ Wang, and KF Kelton. Experimental and molecular dynamics simulation study of structure of liquid and amorphous ni62nb38 alloy. The Journal of chemical physics, 145(20):204505, 2016.

36 MATERIALS SCIENCE↗

"Traffic Control via Connected and Automated Vehicles: An Open-Road Field Experiment with 100 CAVs"

The CIRCLES project aims to reduce instabilities in traffic flow, which are naturally occurring phenomena due to human driving behavior. These "phantom jams" or "stop-and-go waves,"are a significant source of wasted energy. Toward this goal, the CIRCLES project designed a control system referred to as the MegaController by the CIRCLES team, that could be deployed in real traffic. Our field experiment leveraged a heterogeneous fleet of 100 longitudinally-controlled vehicles as Lagrangian traffic actuators, each of which ran a controller with the architecture described in this paper. The MegaController is a hierarchical control architecture, which consists of two main layers. The upper layer is called Speed Planner, and is a centralized optimal control algorithm. It assigns speed targets to the vehicles, conveyed through the LTE cellular network. The lower layer is a control layer, running on each vehicle. It performs local actuation by overriding the stock adaptive cruise controller, using the stock on-board sensors. The Speed Planner ingests live data feeds provided by third parties, as well as data from our own control vehicles, and uses both to perform the speed assignment. The architecture of the speed planner allows for modular use of standard control techniques, such as optimal control, model predictive control, kernel methods and others, including Deep RL, model predictive control and explicit controllers. Depending on the vehicle architecture, all onboard sensing data can be accessed by the local controllers, or only some. Control inputs vary across different automakers, with inputs ranging from torque or acceleration requests for some cars, and electronic selection of ACC set points in others. The proposed architecture allows for the combination of all possible settings proposed above. Most configurations were tested throughout the ramp up to the MegaVandertest.

Lee, Jonathan↗

Causes of and Solutions to Wind Speed Bias in NREL's 2020 Offshore Wind Resource Assessment for the California Pacific Outer Continental Shelf

This report provides the results of a detailed analysis into the causes of high wind speed bias in the 20-year wind resource data set for offshore California the National Renewable Energy Laboratory (NREL) released in 2020, herein called CA20. The data set was developed using the state-of-the-art Weather Research and Forecasting (WRF) model. Notably, no floating lidars were available at the time in offshore California to validate offshore hub-height wind speeds. In late 2020, the Pacific Northwest National Laboratory (PNNL) deployed two floating lidars in the California outer continental shelf (OCS), near the Bureau of Ocean Energy Management (BOEM) call areas of Humboldt and Morro Bay. Using these observations through 2021, NREL found considerable bias in modeled hub-height winds at both locations: up to +2 m/s at Humboldt over a 6-month period, and up to +1 m/s at Morro Bay over a one-year period. Upon the discovery of this bias, the Department of Energy (DOE) and BOEM funded NREL and PNNL to investigate the causes of, impacts of, and solutions to the bias in the CA20 data set. This report summarizes the findings of this research. We first investigated whether different WRF model setups could lead to reduced bias. We found that the choice of planetary boundary layer (PBL) scheme - which controls the vertical turbulent mixing of momentum, heat, and moisture in the lowermost part of the atmosphere - greatly affected hub-height wind speeds in the region. Specifically, switching from the Mellor-Yamada-Nakanishi-Niino (MYNN) scheme used in CA20 (and widely used across a range of operational and research weather models) to the less common Yonsei University (YSU) scheme nearly eliminated the bias at both the Humboldt and Morro Bay lidar locations. The large discrepancy between the MYNN- and YSU-modeled hub-height winds pointed towards the role of atmospheric stability. In general, PBL schemes agree well in conditions of high turbulence and mixing, normally referred to as "unstable" conditions. By contrast, PBL schemes start to diverge in "stable" conditions, where turbulence is low and thermal stratification (i.e., higher temperature air sitting on top of colder air) greatly suppresses vertical mixing. Under such conditions, winds aloft can decouple from surface effects and greatly accelerate, causing high wind speeds at hub-height and frequent low-level jets (LLJs). We determined that these stable conditions are in fact dominant in offshore California. The region is characterized by moderate-to-extreme stable stratification with a LLJ on average around 200 meters above sea-level. To our knowledge, no wind energy area globally has as strongly stable stratification as offshore California. Under these extreme conditions, we determined that the MYNN scheme models higher stability than YSU, resulting in less vertical turbulent mixing than YSU, allowing for the acceleration of hub-height winds, more intense LLJs, and higher-amplitude inertial oscillations. Using surface observations, we found that MYNN overestimates near-surface stability, whereas YSU tends to model stability better. We then considered several short-term case studies to assess additional meteorological drivers of the bias at Humboldt. We found that during synoptic scale northerly flows driven by the North Pacific High and inland thermal low, a coastal warm bias in the MYNN case studies contributes to the modeled wind speed bias by altering the boundary layer thermodynamics via a thermal wind mechanism. Given the strong performance of the YSU-based runs in offshore California, NREL has produced and published an updated version of the CA20 data set with YSU as the PBL scheme. This updated data set is now part of NREL's 2023 National Offshore Wind (NOW-23) data set, which covers all the U.S. offshore waters. The development and final validation of the NOW-23 data set in offshore California is documented in this report.

17 WIND ENERGY↗

Computing the Properties of Matter with Leadership Computing Resources (Closeout Report for DE-SC0018121)

In order to add more capabilities to Halide, we have designed a new framework called Tiramisu and integrated this framework into Halide. Since Tiramisu enables Halide to target heterogeneous architectures, our development efforts have been refocused on Tiramisu. Most high-performance computer systems today are complex and increasingly heterogeneous; they may have CPUs, GPUs and FPGAs. Achieving best performance requires taking full advantage of all these different architectures. To address this issue, we have designed Tiramisu, an optimization framework that enables Halide (and other DSLs) to target heterogeneous architectures. Tiramisu is an optimization framework that takes as input a high level, architecture-independent representation of code and a set of scheduling and data mapping commands that guide code transformation. The input can either be generated by a domain-specific language (DSL) compiler such as Halide or directly written by a programmer. Tiramisu then applies the user-specified code and data-layout transformations and generates an architecture-specific, low-level intermediate representation (IR) that takes advantage of modern architectural features such as multicore parallelism, non-uniform memory (NUMA) hierarchies, clusters, and accelerators like GPUs and FPGAs. We integrated Tiramisu within Halide and implemented a representative set of benchmarks to evaluate this integration. Tiramisu is now open source and is available for public use (http://tiramisu-compiler.org/). A paper about Tiramisu was published, it shows that Tiramisu extends Halide with many new capabilities and that Tiramisu can generate efficient code for multicores, GPUs, FPGAs and distributed heterogeneous systems. The performance of code generated by the Tiramisu backends matches or exceeds hand optimized reference implementations. For example, the multicore backend matches the highly optimized Intel MKL library on many kernels and shows speedups reaching 4x over the original Halide. In addition to making Tiramisu more robust, we have used Tiramisu to implement a set of representative tensor operation for constructing baryon building blocks required for multi baryon contractions in LQCD. In order to implement this code, we needed to generalize Tiramisu in two ways: first we needed to support indirect array accesses, and second, we needed to add support for complex numbers to Tiramisu. The code generated by Tiramisu is 6x faster than the reference code. Our efforts towards an MPI based multi-node version of tiramisu have matured and the resulting code scales well on multiple nodes (tests up to 512 KNL nodes have been undertaken).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Traffic Control via Connected and Automated Vehicles (CAVs): An Open-Road Field Experiment with 100 CAVs

The CIRCLES project aims to reduce instabilities in traffic flow, which are naturally occurring phenomena due to human driving behavior. Also called “phantom jams” or “stop-and-go waves,” these instabilities are a significant source of wasted energy. Toward this goal, the CIRCLES project designed a control system, referred to as the MegaController by the CIRCLES team, that could be deployed in real traffic. Our field experiment, the MegaVanderTest (MVT), leveraged a heterogeneous fleet of 100 longitudinally controlled vehicles as Lagrangian traffic actuators, each of which ran a controller with the architecture described in this article. The MegaController is a hierarchical control architecture that consists of two main layers. The upper layer is called the Speed Planner and is a centralized optimal control algorithm. It assigns speed targets to the vehicles, conveyed through the LTE cellular network. The lower layer is a control layer, running on each vehicle. It performs local actuation by overriding the stock adaptive cruise controller, using the stock onboard sensors. The Speed Planner ingests live data feeds provided by third parties as well as data from our own control vehicles and uses both to perform the speed assignment. The architecture of the Speed Planner allows for the modular use of standard control techniques, such as optimal control, model predictive control (MPC), kernel methods, and others. The architecture of the local controller allows for the flexible implementation of local controllers. Corresponding techniques include deep reinforcement learning (RL), MPC, and explicit controllers. Depending on the vehicle architecture, all onboard sensing data can be accessed by the local controllers or only some. Likewise, control inputs vary across different automakers, with inputs ranging from torque or acceleration requests for some cars to electronic selection of adaptive cruise control (ACC) setpoints in others. The proposed architecture technically allows for the combination of all possible settings proposed previously, that is {Speed Planner algorithms} × {local Vehicle Controller algorithms} × {full or partial sensing} × {torque or speed control}. As a result, most configurations were tested throughout the ramp up to the MegaVandertest (MVT).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗