Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “open data format”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

The Need for a New LLNL Pulsed Sphere Neutron Leakage Spectra Series

Here, it is shown that spectra measured as part of the Lawrence Livermore National Laboratory Pulsed Sphere (LPS) program offer decisive information to locate formatting or physics issues in nuclear data of key interest for fusion reactor simulations. However, experiments from this measurement series are not benchmarks. For instance, their uncertainties are incomplete. There are also many open questions—e.g., on the setup, the detector response, and whether LPS are accurately modeled—that cannot be answered anymore given the limited documentation and that many of the experimenters are no longer actively working. This limited knowledge has implications when one tries to adjust nuclear data to LPS spectra. Usually, one adjusts to benchmarks representing an application with the hope to get more precise nuclear data for the application of interest where differential data might be scarce and/or to reduce nuclear data uncertainties in the application simulations. However, it is demonstrated that adjustment with LPS spectra without accounting for missing uncertainties and modeling potential biases in the experimental data leads to adjusted data that are highly unphysical. That means adjusted data differ significantly from evaluated data based on information from differential experiments; also, application quantities predicted with the adjusted data deviate distinctly from experimental ones. While we can approximate our limited knowledge on these experiments with Gaussian processes in the adjustment process, this modeling of bias is arbitrary rather than based on a physics explanation, calling into doubt the validity of resulting adjusted data. Thus, we discuss here the need for a new measurement series, learning from the strengths and weaknesses of the LPS program, to yield decisive and well-benchmarked integral experiments to support fusion reactor research.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Wind Turbine Sound Setbacks and Supply Curves: Ordinances and Extrapolated Trends, 110 Hub Height, 130 Rotor Diameter

This dataset provides a comprehensive set of wind turbine sound setbacks from every residential structure in the contiguous United States (CONUS). A sound setback is defined as the minimum required distance between a residential structure and a hypothetical turbine installation site to ensure that modeled sound levels received at the residence do not exceed local sound ordinances, which are commonly expressed in A-weighted decibels (dBA). Therefore, sound setbacks are a local spatial assessment combining multiple factors, including the sound pressure curve as a function of the observer location (distance and direction) relative to the turbine, local sound regulations, and the geographical distribution of residential structures. The dataset is organized into multiple scenario-based products, detailed as follows: 1. Existing and extrapolated sound setbacks. An existing scenario characterizes sound setbacks only in states or counties that have implemented sound regulations as of 2022. The extrapolated scenarios extend a constant sound threshold to counties that lack explicit sound regulations, with thresholds ranging from 35 to 60 dBA, in 5-dBA increments reflecting the variation observed in current sound ordinances. 2. Sound setbacks in directional and worst scenarios. The directional scenario accounts for the distance and orientation of residential structures relative to a hypothetical turbine location, utilizing the turbine's sound emissions in that specific direction. In contrast, the worst scenario takes loudest sound level at each distance step from the turbine, irrespective of directional considerations, which aligns with current industry practice. 3. Supply curves for Open and Reference Access scenarios. This dataset includes supply curves generated by the reV model, which integrates each of the above sound setbacks into both Open and Reference siting scenarios. In addition, two Open and Reference baselines scenarios were included which do not consider sound setbacks for comparative analysis. All sound setback data are stored in TIF files, with partial maps of the data provided in PNG format. The values in the sound setback raster range from 0 to 1, representing the fraction of developable land within a 90 meter by 90 meter pixel due to sound ordinances. A value of 0 indicates areas where wind energy development is prohibited, while a value of 1 signifies areas fully permissible. The wind turbine parameters used in the sound modeling are based on the land-based turbine from International Energy Agency (IEA), featuring a rated electrical power of 3.4 MW, a rotor diameter of 130 meters, and a hub height of 110 meters. The atmospheric conditions, including wind speed/direction, turbulence, air temperature, relative humidity, and air pressure, that drive the sound generation are obtained from the WIND Toolkit dataset.

Array↗

An optimized FM-index library for nucleotide and amino acid search

Abstract Background Pattern matching is a key step in a variety of biological sequence analysis pipelines. The FM-index is a compressed data structure for pattern matching, with search run time that is independent of the length of the database text. Implementation of the FM-index is reasonably complicated, so that increased adoption will be aided by the availability of a fast and flexible FM-index library. Results We present AvxWindowedFMindex (AWFM-index), a lightweight, open-source, thread-parallel FM-index library written in C that is optimized for indexing nucleotide and amino acid sequences. AWFM-index introduces a new approach to storing FM-index data in a strided bit-vector format that enables extremely efficient computation of the FM-index occurrence function via AVX2 bitwise instructions, and combines this with optional on-disk storage of the index’s suffix array and a cache-efficient lookup table for partial k-mer searches. The AWFM-index performs exact match count and locate queries faster than SeqAn3’s FM-index implementation across a range of comparable memory footprints. When optimized for speed, AWFM-index is $$\sim $$ ∼ 2–4x faster than SeqAn3 for nucleotide search, and $$\sim $$ ∼ 2–6x faster for amino acid search; it is also $$\sim $$ ∼ 4x faster with similar memory footprint when storing the suffix array in on-disk SSD storage. Conclusions AWFM-index is easy to incorporate into bioinformatics software, offers run-time performance parameterization, and provides clients with FM-index functionality at both a high-level (count or locate all instances of a query string) and low-level (step-wise control of the FM-index backward-search process). The open-source library is available for download at https://github.com/TravisWheelerLab/AvxWindowFmIndex.

59 BASIC BIOLOGICAL SCIENCES↗

Evidence for Topological Protection Derived from Six-Flux Composite Fermions

The composite fermion theory opened a new chapter in understanding many-body correlations through the formation of emergent particles. The formation of two-flux and four-flux composite fermions is well established. While there are limited data linked to the formation of six-flux composite fermions, topological protection associated with them is conspicuously lacking. Here we report evidence for the formation of a quantized and gapped fractional quantum Hall state at the filling factor ν = 9/11, which we associate with the formation of six-flux composite fermions. Our result provides evidence for the most intricate composite fermion with six fluxes and expands the already diverse family of highly correlated topological phases with a new member that cannot be characterized by correlations present in other known members. Our observations pave the way towards the study of higher order correlations in the fractional quantum Hall regime.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Multidimensional Modeling of Mixture Formation in a Hydrogen-Fueled Heavy-Duty Optical Engine With Direct Injection

Hydrogen (H 2 ), as a carbon-free fuel, is considered as one of the most promising solutions to reduce the carbon footprint of hard-to-decarbonize energy and transportation sectors. As such, hydrogen-fueled internal combustion engines (H 2 ICEs) have recently been receiving increasing attention, particularly in applications such as on-road/off-road heavy-duty transport and combined heat and power. The direct injection (DI) of gaseous hydrogen into the combustion chamber offers great potential for achieving high power density and high engine efficiency, while mitigating the risk of backfire and reducing pre-ignition. However, the numerical simulation of H 2 DI system remains a formidable challenge associated with the high computational cost of reproducing compressible supersonic flow and shocks in narrow injector passages and in near-nozzle regions. In general, there is a lack of well-established and validated practices for the modeling of high-pressure H 2 DI in large-bore engines. Here, to this end, this study focuses on computational fluid dynamics (CFD) modeling of the mixture formation process in a heavy-duty optical engine employing a medium-pressure H 2 DI system. Both large eddy simulations (LES) and Reynolds Averaged Navier–Stokes (RANS) simulations are performed and evaluated against optical data. Gaseous hydrogen is injected into the combustion chamber via a centrally located outward opening hollow-cone injector at a pressure of 40 bar. Simulations are carried out for two injection timings, namely, −120 and −60 °CA. The numerical predictions for H 2 distribution in different horizontal and vertical planes during the compression stroke are systematically compared against optical data obtained through planar laser-induced fluorescence (PLIF) measurements. Overall, the LES approach using the Dynamic Structure model is found to have good predictive capabilities for the early jet penetration in terms of length and shape, as well as the later H 2 distributions. However, the unsteady RANS approach with the renormalization group $k - ϵ$ model, which is widely used by industry to model heavy-duty ICEs, significantly underpredicts the H 2 mixing, even at similar mesh resolution to that used in LES. These results indicate that there is a need for the improvement of mixing submodels within the RANS approach when applied to H 2 DI simulations.

LES↗

Separation of CO 2 from Flue Gas and Potential for Geologic Sequestration

The objectives of this study were to review various methods reported in the literature for the separation and geologic sequestration of carbon dioxide and evaluate the potential of TVA fossil fuel-burning plant locations for onsite geologic sequestration of CO 2 from stack emissions. Several conventional and nonconventional technologies for the separation of CO 2 from flue gas, including absorption, adsorption, cryogenic distillation, membranes, hydrate formation and dissociation, and ammonia carbonation, have been reviewed in terms of separation mechanisms, flow diagrams, and costs. Most of the technologies that have been reviewed are still at the research and development stage. Critical information needed to assess and compare these technologies is still lacking. In addition, information on some of the technologies that have been tested at a pilot or industrial scale has not been fully disclosed in the open literature. Because of this lack of data, it is difficult to make a critical assessment of each of the separation technologies. Based on limited information, it was concluded that the most promising methods are membrane separation and the Mitsubishi process for chemical absorption. Both processes involve separating CO 2 at high temperature, minimizing the cost for cooling prior to separation. Physical and chemical geologic formations of CO 2 were also reviewed. It was concluded that due to the geologic time scale of CO 2 sequestration periods, relatively safe conditions, general proximity to CO 2 sources, and extensive knowledge of underground conditions, sequestration of CO 2 in underground aquifers and coal beds is a very promising method of mitigating greenhouse gas emissions. The cost is predicted to be relatively low and the suitable sites are numerous for this application, with many of these sites located close to the plants.

20 FOSSIL-FUELED POWER PLANTS↗

NEON AOP Survey of Upper East River CO Watersheds: Waveform LiDAR Binary Data

The waveform Light Detection and Ranging (LiDAR) data in this package were generated through a National Ecological Observatory Network Airborne Observation Platform (NEON AOP) acquisition over watersheds of interest surrounding Crested Butte, Colorado. The remote sensing imagery acquired by the NEON AOP supports an interdisciplinary project on hydrology, biogeochemistry, and ecosystem functioning in a snow-dominated headwater environment. These waveform LiDAR data enable spatially continuous estimation of vegetation structure parameters to facilitate analyses of the major environmental drivers of structural and compositional variability. The package contains 97 compressed file archives in 7-zip (.7z) format, each corresponding to one acquisition flightpath. Within each .7z archive is a set of constituent files describing properties of the LiDAR waveforms, such as return intensity, geolocation, outgoing pulse and other behavior of the sensor and signals. Once downloaded, the files must first be unzipped using the widely distributed command-line software utility 7z, using the command '7z x \[filename\].7z \[target_directory\]'. All files within the .7z archives can be opened in IDL, MatLab, or the open-source R statistical computing environment. Further details about the data package are in the attached user guide (neon_aop_crbu_waveformlidar_userguide.pdf). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Data-Driven Buy Clean: Decarbonization and Beyond

This report was compiled to provide recommendations on the availability of public background data from the U.S. Federal life cycle assessment (LCA) Data Commons to be conformant with the Association for Life Cycle Assessment (ACLCA) 2022 Product Category Rule (PCR) Open Standard to build technical tools that can assist industry in creating more comparable Type II Environmental Product Declarations (EPDs) for Federal Buy Clean and sustainability initiatives. The Federal LCA Commons is not only a public data source but also a consistently structured, self-referencing mega-repository for data developed by federal agency experts (in agency repositories) and by academia, nonprofit organizations, and industry (via the US Life Cycle Inventory Database). The Federal LCA Commons Technical Working Group is continuously improving the standardization of data documentation, formatting, and nomenclature to ensure lossless data loading and accurate data representation. This report and appendixes include the following: 1) An introduction to data-driven Buy Clean and decarbonization initiatives at the federal level; 2) The current status and associated challenges with LCA data and EPD standards and comparability; 3) Opportunities for the Federal LCA Commons to support conformance with the ACLCA 2022 PCR Open Standard and provide resources to implement the Federal Sustainability Plan, Buy Clean Program, and Inflation Reduction Act (IRA) sustainability goals and objectives. To date, the Federal LCA Commons is the result of coordinated work by National Renewable Energy Laboratory (NREL), the U.S. Department of Agriculture (USDA), the Environmental Protection Agency (EPA), the National Energy Technology Laboratory (NETL), the Argonne National Laboratory (ANL), the U.S. Army Corps of Engineers (USACE), the Federal Highway Administration (FHWA), the U.S. Forest Service (USFS), the Federal Aviation Administration (FAA), the Department of Defense (DoD) and the National Institute of Standards and Technologies (NIST). The Federal LCA Commons will continue to combine databases from the collaborating agencies while remaining a public resource. There are several initiatives among the collaborating agencies to expand the Federal LCA Commons and dedicated federal funding and resources could accelerate and strengthen these initiatives.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Flexible, integrated modeling of tokamak stability, transport, equilibrium, and pedestal physics

The STEP (Stability, Transport, Equilibrium, and Pedestal) integrated-modeling tool has been developed in OMFIT to predict stable, tokamak equilibria self-consistently with core-transport and pedestal calculations. STEP couples theory-based codes to integrate a variety of physics, including magnetohydrodynamic stability, transport, equilibrium, pedestal formation, and current-drive, heating, and fueling. The input/output of each code is interfaced with a centralized ITER-Integrated Modelling & Analysis Suite data structure, allowing codes to be run in any order and enabling open-loop, feedback, and optimization workflows. This paradigm simplifies the integration of new codes, making STEP highly extensible. STEP has been verified against a published benchmark of six different integrated models. Core-pedestal calculations with STEP have been successfully validated against individual DIII-D H-mode discharges and across more than 500 discharges of the H98,y2 database, with a mean error in confinement time from experiment less than 19%. STEP has also reproduced results in less conventional DIII-D scenarios, including negative-central-shear and negative-triangularity plasmas. Predictive STEP modeling has been used to assess performance in several tokamak reactors. Simulations of a high-field, large-aspect-ratio reactor show significantly lower fusion power than predicted by a zero-dimensional study, demonstrating the limitations of scaling-law extrapolations. STEP predictions have found promising scenarios for an EXhaust and Confinement Integration Tokamak Experiment, including a high-pressure, 80%-bootstrap-fraction plasma. ITER modeling with STEP has shown that pellet fueling enhances fusion gain in both the baseline and advanced-inductive scenarios. Finally, STEP predictions for the SPARC baseline scenario are in good agreement with published results from the physics basis.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

OpenMDlr: parallel, open-source tools for general protein structure modeling and refinement from pairwise distances

Easy-to-use, open-source, general-purpose programs for modeling a protein structure from inter-atomic distances are needed for modeling from experimental data and refinement of predicted protein structures. OpenMDlr is an open-source Python package for modeling protein structures from pairwise distances between any atoms, and optionally, dihedral angles. Finally, we provide a user-friendly input format for harnessing modern biomolecular force fields in an easy-to-install package that can efficiently make use of multiple compute cores.

59 BASIC BIOLOGICAL SCIENCES↗

Symmetry-mode analysis for local structure investigations using pair distribution function data

Symmetry-adapted distortion modes provide a natural way of describing distorted structures derived from higher-symmetry parent phases. Structural refinements using symmetry-mode amplitudes as fit variables have been used for at least ten years in Rietveld refinements of the average crystal structure from diffraction data; more recently, this approach has also been used for investigations of the local structure using real-space pair distribution function (PDF) data. Here, the value of performing symmetry-mode fits to PDF data is further demonstrated through the successful application of this method to two topical materials: TiSe2, where a subtle but long-range structural distortion driven by the formation of a charge-density wave is detected, and MnTe, where a large but highly localized structural distortion is characterized in terms of symmetry-lowering displacements of the Te atoms. Here, the analysis is performed using fully open-source code within the DiffPy framework via two packages developed for this work: isopydistort, which provides a scriptable interface to the ISODISTORT web application for group theoretical calculations, and isopytools, which converts the ISODISTORT output into a DiffPy-compatible format for subsequent fitting and analysis. These developments expand the potential impact of symmetry-adapted PDF analysis by enabling high-throughput analysis and removing the need for any commercial software.

36 MATERIALS SCIENCE↗

Decomprolute is a benchmarking platform designed for multiomics-based tumor deconvolution

Tumor deconvolution is a reliable way to disentangle the diverse cell types that comprise solid tumors. To date, however, both the algorithms developed to deconvolve tumor samples, and the gold standard datasets used to assess the algorithms are geared toward the analysis of gene expression (e.g., RNA-seq) rather than protein levels in tumor cells. While gene expression is less expensive to measure, protein levels provide a more accurate view of immune markers. To facilitate the development as well as improve the reproducibility and reusability of multi-omic deconvolution algorithms, we introduce Decomprolute, a Common Workflow Language framework that leverages containerization to compare tumor deconvolution algorithms across multiomic data sets. Decomprolute incorporates the large-scale multiomic data sets produced by the Clinical Proteomic Tumor Analysis Consortium (CPTAC), which include matched mRNA expression and proteomic data from thousands of tumors across multiple cancer types to build a fully open-source, containerized proteogenomic tumor deconvolution benchmarking platform. The platform consists of modular architecture and it comes with well-defined input and output formats at each module. As a result, it is robust and extendable easily with additional algorithms or analyses. The platform is available for access and use at http://pnnl-compbio.github.io/decomprolute.

60 APPLIED LIFE SCIENCES↗

Aerial Captured Data and Processed Models in Beaumont-Port Arthur Region in Feb and Oct, 2023

Our Co-design team is from the University of Texas, working on a Department of Energy-funded project focused on the Beaumont-Port Arthur area. As part of this project, we will be developing climate-resilient design solutions for areas of the region. More on www.caee.utexas.edu.We used a DJI Mavic 2 Pro to capture aerial photos in Beaumont-Port Arthur, TX, in February 2023, including:I. Beaumont Soccer ClubII. Corps’ Port Arthur Resident OfficeIII. Halbouty Pump Station comprises its vicinityIV. Lamar University (Including Exxon Power Plants close to Lamar Univ.)V. MLK Boulevard for aerial images of the industry and the ship channelVI. Salt Water Barrier (include some aerial images about the Big Thicket)Aerial photos taken were through DroneDeploy autonomous flight, and models were processed through the DroneDeploy engine as well. All aerial photos are in .JPG format and contained in zipped files for each location.The processed data package including 3D models, geospatial data, mappings, point clouds, and the animation video of Halbouty Pump Station has various file types:- The Adobe Suite gives you great software to open .Tif files.- You can use LASUtility (Windows), ESRI ArcGIS Pro (Windows), or Blaze3D (Windows, Linux) to open a LAS file and view the data it contains.- Open an .OBJ file with a large number of free and commercial applications. Some examples include Microsoft 3D Builder, Apple Preview, Blender, and Autodesk.- You may use ArcGIS, Merkaartor, Blender (with the Google Earth Importer plug-in), Global Mapper, and Marble to open .KML files.- The .tfw world file is a text file used to georeference the GeoTIFF raster images, like the orthomosaic and the DSM. You need suitable software like ArcView to open a .TFW file.This dataset provides researchers with sufficient geometric data and the status quo of the land surface at the locations mentioned above. This dataset could streamline researchers' decision-making processes and enhance the design as well.In October 2023, we had our follow-up data collection, including:I. Beaumont Soccer ClubII. Shipping and Receiving Center at Lamar UniversityAfter the aerial collection, we obtained aerial photos of those two locations mentioned above, as well as processed data (such as point clouds and models).

2D mapping↗

Formation of Point Shocks for 3D Compressible Euler

We consider the 3D isentropic compressible Euler equations with the ideal gas law. We provide a constructive proof of the formation of the first point shock from smooth initial datum of finite energy, with no vacuum regions, with nontrivial vorticity present at the shock, and under no symmetry assumptions . We prove that for an open set of Sobolev‐class initial data that are a small L ∞ perturbation of a constant state, there exist smooth solutions to the Euler equations which form a generic stable shock in finite time. The blowup time and location can be explicitly computed, and solutions at the blowup time are smooth except for a single point , where they are of cusp‐type with Hölder C 1/3 regularity. Our proof is based on the use of modulated self‐similar variables that are used to enforce a number of constraints on the blowup profile, necessary to establish global existence and asymptotic stability in self‐similar variables. © 2022 Wiley Periodicals LLC.

Mathematics↗

MTUQ: a framework for estimating moment tensors, point forces, and their uncertainties

SUMMARY We introduce MTUQ, an open-source Python package for seismic source estimation and uncertainty quantification, emphasizing flexibility and operational scalability. MTUQ provides MPI-parallelized grid search and global optimization capabilities, compatibility with 1-D and 3-D Green’s function database formats, customizable data processing, C-accelerated waveform and first-motion polarity misfit functions, and utilities for plotting seismic waveforms and visualizing misfit and likelihood surfaces. Applicability to a range of full- and constrained-moment tensor, point force, and centroid inversion problems is possible via a documented application programming interface, accompanied by example scripts and integration tests. We demonstrate the software using three different types of seismic events: (1) a 2009 intraslab earthquake near Anchorage, Alaska; (2) an episode of the 2021 Barry Arm landslide in Alaska; and (3) the 2017 Democratic People’s Republic of Korea underground nuclear test. With these events, we illustrate the well-known complementary character of body waves, surface waves, and polarities for constraining source parameters. We also convey the distinct misfit patterns that arise from each individual data type, the importance of uncertainty quantification for detecting multimodal or otherwise poorly constrained solutions, and the software’s flexible, modular design.

58 GEOSCIENCES↗

Topography and canopy cover influence soil organic carbon composition and distribution across a forested hillslope in the discontinuous permafrost zone

This dataset contains data used for the paper "Topography and canopy cover influence soil organic carbon composition and distribution across a forested hillslope in the discontinuous permafrost zone". The Related References field will be updated with a full citation when available. Topography and canopy cover influence ground temperature in warming permafrost landscapes, yet soil temperature heterogeneity introduced by meso-topographic slope positions, microtopographic differences in vegetation cover, and the subsequent impact of contrasting temperature conditions to soil organic carbon (SOC) dynamics are understudied. Buffering of permafrost-affected soils against warming air temperatures in boreal forests can reflect surface soil characteristics (e.g., thickness of organic material) as well as the degree and type of canopy cover (e.g., open cover vs closed cover). Both landscape and soil properties interact to determine meso- and micro-scale heterogeneity of ground warming. We sampled a hillslope catena transect in a discontinuous permafrost zone near Fairbanks, Alaska to test the small-scale (1 to 3 meter) impacts of slope position and cover type on soil organic matter composition. Mineral active layer samples were collected from backslope, low backslope, and footslope positions at depths spanning 19 to 60 cm. We examined soil mineralogical composition, soil moisture, total carbon and nitrogen content, and organic mat thickness in conjunction with an assessment of SOC composition using Fourier-transform ion Cyclotron Resonance Mass Spectrometry (FT-ICR-MS). Soils in the footslope position had a higher relative contribution of lignin-like compounds while backslope soils had more aliphatic and condensed aromatic compounds as determined by FT-ICR-MS. The effect of open versus closed tree canopy cover varied with slope position. On the backslope, we found higher oxidation of molecules under open cover compared with closed cover, indicating an effect of warmer soil temperature on decomposition. Little to no effect of canopy was observed for soils at the footslope position, which we attributed, in part, to the strong impact of soil moisture content in SOC dynamics in the water-gathering footslope position. The thin organic mat under open cover on the backslope position may have contributed to differences in soil temperature and thus SOC oxidation under open and closed canopy. Here, the thinner organic mat did not appear to buffer the underlying soil against warm season air temperatures and thus increased SOC decomposition as indicated by higher oxidation of SOC molecules and a lower contribution of simple molecules under open cover compared with the closed canopy sites. Our findings suggest that the role of canopy cover in SOC dynamics varies as a function of landscape position and soil properties, namely organic mat thickness and soil moisture. Condition-specific heterogeneity of SOC composition under open and closed canopy cover highlights the protective effect of canopy cover for soils on backslope positions. This dataset contains a compressed (.zip) archive of the data and R scripts used for this manuscript. The dataset includes files in .csv format, which can be accessed and processed using MS Excel or R. This archive can also be accessed on GitHub at https://github.com/Erin-Rooney/Y1_fairbanks (DOI: 10.5281/zenodo.8071247).

54 ENVIRONMENTAL SCIENCES↗

Machine learning model inputs, outputs, and scripts associated with “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions” (Malhotra et al., in prep). This effort was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the contiguous United States (CONUS). New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Associated sediment and water geochemistry and in situ sensor data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775. This data package is associated with two GitHub repositories found at https://github.com/parallelworks/dynamic-learning-rivers and https://github.com/WHONDRS-Hub/ICON-ModEx_Open_Manuscript. In addition to this readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This data package consists of two main folders (1) dynamic-learning-rivers and (2) ICON-ModEx_Open_Manuscript which contain snapshots of the associated GitHub repositories. The input data, output data, and machine learning models used to guide sampling locations are within dynamic-learning-rivers. The folder is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning (ML) models trained on the data in “input_data”; (3) “examples” contains files for direct experimentation with the machine learning model, including scripts for setting up “hindcast” run; (4) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; and (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please see the top-level README.md in the GitHub repository for more details on the automation. The scripts and data used to create figures in the manuscript are within ICON-ModEx_Open_Manuscript. The folder is organized into four folders which contain the scripts, data, and pdf for each figure. Within the “fig-model-score-evolution” folder, there is a folder called “intermediate_branch_data” which contains some intermediate files pulled from dynamic-learning-rivers and reorganized to easily integrate into the workflows. NOTE: THIS FOLDER INCLUDES THE FILES AT THE POINT OF PAPER SUBMISSION. IT WILL BE UPDATED ONCE THE PAPER IS ACCEPTED WITH ANY REVISIONS AND WILL INCLUDE A DD/FLMD AT THAT POINT. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

Bond-centric modular design of protein assemblies

Directional interactions that generate regular coordination geometries are a powerful means of guiding molecular and colloidal self-assembly, but implementing such high-level interactions with proteins remains challenging due to their complex shapes and intricate interface properties. Here we describe a modular approach to protein nanomaterial design inspired by the rich chemical diversity that can be generated from the small number of atomic valencies. We design protein building blocks using deep learning-based generative tools, incorporating regular coordination geometries and tailorable bonding interactions that enable the assembly of diverse closed and open architectures guided by simple geometric principles. Experimental characterization confirms the successful formation of more than 20 multicomponent polyhedral protein cages, two-dimensional arrays and three-dimensional protein lattices, with a high (10%–50%) success rate and electron microscopy data closely matching the corresponding design models. Due to modularity, individual building blocks can assemble with different partners to generate distinct regular assemblies, resulting in an economy of parts and enabling the construction of reconfigurable networks for designer nanomaterials.

Biomaterials – proteins↗