Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Piston geometry and stroke optimization for high efficiency propane spark ignition engines

Propane has unique properties and offers interesting characteristics for high-efficiency spark ignition engines. Its high volatility reduces or completely eliminates fuel-wall wetting and facilitates fuel air mixing. Furthermore, propane has a research octane number of 112 and a high octane sensitivity of 15. Finally, its laminar flame speed is on the same order as that of conventional gasoline, and it exhibits high dilution tolerance. Modern spark ignition internal combustion engines rely on fast combustion rates and high dilution to achieve high brake thermal efficiencies. To accomplish this, high stroke-to-bore ratios and high geometric compression ratios have been used in new engine designs. Therefore, propane’s relatively high laminar flame speeds, high knock resistance, and dilution tolerance make it an excellent candidate fuel for modern spark ignition engines. The objective of this work is to co-optimize the piston geometry and the engine stroke to maximize the efficiency of a spark-ignition engine fueled with propane. 3D computational fluid dynamics (CFD) simulations employing the extended coherent flamelet model were used to study the parametric effects of piston shape and stroke length. A piston geometry based on high performing pistons was parameterized using four controlling parameters. The piston geometry and engine stroke design space was explored using deterministic and quasi-random sampling techniques. In conclusion, a Gaussian process regression model was built using the simulation data to explain the results observed.

33 ADVANCED PROPULSION SYSTEMS↗

Development of physics-consistent conditional diffusion model to overcome data scarcity in critical heat flux

Deep generative modeling provides a powerful pathway to overcome data scarcity in energy-related applications where experimental data are often limited. By learning the underlying probability distribution of the training dataset, deep generative models, such as the diffusion model, can generate high-fidelity synthetic samples that statistically resemble the training data. Such synthetic data generation can significantly enrich the size and diversity of the available training data, and more importantly, improve the robustness of downstream machine learning models in predictive tasks. The objective of this paper is to investigate the effectiveness of diffusion models for overcoming data scarcity in nuclear energy applications. By leveraging a public dataset on critical heat flux which covers a wide range of commercial nuclear reactor operational conditions, we developed a diffusion model that can generate an arbitrary amount of synthetic samples. Since a vanilla diffusion model can only generate samples randomly, we also developed a conditional diffusion model capable of generating targeted critical heat flux data under user-specified thermal-hydraulic conditions. The performance of the diffusion model was evaluated based on its ability to capture empirical feature distributions and pair-wise correlations, as well as to maintain physical consistency. The results showed that both the diffusion model and conditional diffusion model can successfully generate realistic and physics-consistent critical heat flux data. Furthermore, uncertainty quantification results demonstrate that the conditional diffusion model is highly effective in augmenting critical heat flux data while maintaining acceptable levels of uncertainty.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Multi-stage heat release of multi-component fuels: Insights and implications for advanced engine operation

Multi-stage heat release (MSHR) is a unique phenomenon typically seen in lean/diluted and low- to intermediate-temperature combustion. Despite its relevance to advanced engine operation, the MSHR of multi-component fuels has barely been quantified. This study aims to characterize the MSHR of multi-component fuels in a rapid compression machine (RCM) at conditions representative of advanced combustion engines. New experimental data are first reported in the RCM at an equivalence ratio of 0.4, pressure of 60 bar and temperatures from 702 to 795 K for a research grade, multi-component gasoline surrogate, termed PACE-20. Here, experiments confirm the existence of MSHR for PACE-20 at all temperatures, and reveal the strong inhibiting effect of temperature on MSHR, where increasing temperature inhibits MSHR by suppressing first stage heat release and promoting second and third stage heat release. A detailed chemical kinetic model is also adopted to model the experiments, with good agreement observed. The response of MSHR characteristics to changes in different engine operating parameters (i.e., temperature, pressure, equivalence ratio and CO 2 dilution level) and fuel compositions (i.e., mole fraction of n-pentane, n-heptane, isooctane, cyclopentane, 1-hexene, ethanol and aromatics in PACE-20) is further evaluated via statistical analysis coupling quasi-random sampling, extensive computer experiments and high-dimensional model representation. The change in MSHR is characterized through 8 quantities of interest (QOIs), i.e., duration and extent of each heat release stage, and their standard deviations. The analysis highlights the dependence of the QOIs on the individual parameters as well as their interplays, with temperature exhibiting the strongest impact among all parameters. Furthermore, it is demonstrated that if MSHR is to be facilitated and combustion phasing is to be extended to enable engine operation at higher compression ratios, it is recommended to operate the engine at low intake temperatures, boosted intake pressures and high CO 2 dilutions (i.e., high EGR levels) with gasoline fuels containing more n-pentane and less aromatics.

42 ENGINEERING↗

A probabilistic creep model incorporating test condition, initial damage, and material property uncertainty

Uncertainty is prevalent in the creep resistance of alloys, where at elevated temperature and low pressure, rupture can range across logarithmic decades. In this study, a probabilistic continuum-damage-mechanics (CDM)-based model is derived to capture the uncertainty of creep resistance. To meet this objective, creep data for alloy 304 Stainless Steel is gathered. A constitutive model, “Sinh”, is calibrated deterministically to determine the statistical variability of the material properties. Three sources of uncertainty are injected into the model: test condition (stress and temperature), initial damage, and material properties. Probabilistic simulations are carried out by (a) calibrating probability distribution functions (pdfs) for each source of uncertainty (b) randomly sampling the pdfs using Monte Carlo methods and (c) executing simulations to replicate the uncertain creep behavior. A sensitivity analysis is performed to evaluate the relative effect of each source of uncertainty. In full probabilistic simulations, the cumulative uncertainty of creep behavior is evaluated. The probabilistic model accurately predicts the creep deformation and rupture of the available experiments. The probabilistic model is validated for interpolation but lacks extrapolation ability. Several future works are proposed to further improve the model.

36 MATERIALS SCIENCE↗

Detection of surface water temperature variations of Mongolian lakes benefiting from the spatially and temporally gap-filled MODIS data

Lakes provide critical water resources for human activities and ecosystems, particularly in the Mongolian Plateau (MP), which is characterized by a dry climate and a harsh environment. As a region that is sensitive to anthropogenic warming, tracking lake surface water temperature (LSWT) changes in Mongolian lakes is crucial for understanding the consequences of a warming climate on lake ecosystems. However, the long-term monitoring of LSWT is restricted by the spatiotemporal gaps in the raw imagery of remote sensing-based land surface temperature (LST), e.g., the commonly used Moderate Resolution Imaging Spectroradiometer (MODIS) LST products. This study applied an improved gap-filling method by utilizing the discrete cosine transform-based penalized least squares (DCT-PLS) strategy in the spatial domain combined with the linear interpolation (LI) algorithm in the temporal domain. The method was applied to fill gaps in the LSWT imagery of 12 representative lakes across MP. The randomly sampled high-quality MODIS LSWT values in the spatial and temporal domains were excavated as false data gaps and considered “virtual true” validation datasets. The spatial validation results showed that the estimated LSWT for all the lake cases were comparable with the “virtual true” LSWT values, with the average values of the coefficient of determination, mean absolute error, mean square error, and root mean square error being 0.98, 0.38 °C, 0.45 °C, and 0.59 °C, respectively. Meanwhile, the error of nighttime LSWT results was relatively lower than that of daytime LSWT. For temporal interpolation validation, the LI algorithm exhibited relatively better performance and could more objectively indicate the variation in LSWT. Benefiting from the spatially and temporally well-constrained data, we analyzed the interannual and intra-annual change characteristics of the LSWTs of the 12 lakes. The long-term variations of annual and seasonal mean LSWTs in the 12 selected lakes exhibited no evident trends in 2000–2020, while presented apparent interannual fluctuations. The slight changes in the average LSWTs of the 12 selected lakes were in excellent synchronization with the surrounding LST derived from the reanalysis datasets, confirming the widely reported phenomenon of “global warming hiatus” that occurred in the early 21st century. This study improves the understanding of the LSWT variations in Mongolian lakes in response to global climate change. It has the potential to provide an effective approach for monitoring LSWT changes in other large-scale studies.

54 ENVIRONMENTAL SCIENCES↗

Storylines for the 1997 New Year’s Flood: The role of watershed antecedent conditions and future warming in shaping discharge in the Truckee River watershed

The 1997 New Year’s flood was among the most devastating floods in the Truckee River watershed located in western Nevada. This event resulted from complex interactions of flood drivers, such as extreme precipitation, wet antecedent watershed conditions, warm temperatures and rapid snowmelt. We leveraged simulated forcings from the regionally refined mesh capabilities of the Energy Exascale Earth System Model (RRM-E3SM) and a process-based hydrological model to recreate the 1997 New Year’s flood for the Truckee River watershed across four climate warming levels ranging from the current temperatures to + 4° C. For each scenario, we conducted ensemble simulations with the same forcing but with 100 different seasonal watershed antecedent conditions, which were randomly sampled from long-term hydrological simulations. The results show that the 1997 New Year’s flood can be reproduced or exceeded consistently only when the antecedent watershed conditions are wet, specifically when streamflows are above the 75th percentile of the climatological value. There is negligible change in ensemble mean peakflows for Truckee River near Reno; however, there are increases of 18% and 14% under the warming levels of + 3° C and + 4° C, respectively. The increases in peakflows under future climate warming are attributed to wetter antecedent watershed conditions and enhanced snowmelt. Furthermore, the largest increases in peakflows occur at small, high-elevation headwater basins along the Sierra Nevada crest. This study highlights that changes in extreme flood events will result from the complex interplay of multiple flood drivers. It also demonstrates the potential of storyline approaches to analyze future realizations of these extreme events under different climate scenarios.

Climate change impact study↗

Material Interactions in Severe Accidents – Benchmarking the MELCOR V2.2 Eutectics Model for a BWR-3 Mark-I Station Blackout: Part II – Uncertainty Analysis

Single case comparisons between severe accident simulations can provide detailed insights into severe accident model behavior, however, they cannot offer insights into model uncertainty, sensitivity to uncertain parameters, or underlying model biases.Here in this analysis, the single case benchmark comparison of the MELCOR material interaction models for a station blackout (SBO) scenario of a boiling water reactor (BWR) using representative Fukushima Daiichi Unit 1 boundary conditions is expanded to include an uncertainty analysis. As part of this uncertainty analysis, 1200 simulations are performed for each material interaction model (2400 total), with random sampling of 14 uncertain MELCOR input parameters. Input parameters are selected for their impact on models representing core degradation processes. These include candling, fuel rod failure, debris quenching and dryout. The analysis performed here is not a traditional “best-estimate” uncertainty analysis that uses best-estimate parameters or identifies best-estimate figure of merit distributions. Instead, it is an exploratory uncertainty analysis that identifies and interrogates underlying model form biases of the two material interaction models (eutectics and interactive materials models). Uniform distributions are applied to all uncertain parameters to ensure coverage of the model parameter uncertainty space. Key findings from this study include underlying model form biases exhibited by material interaction models, and notable differences in accident progression outcomes between the material interaction models. This uncertainty study extends and confirms the conclusions from the first part of this study, which compared the impact of material interaction modeling on simulation of a short-term station blackout scenario with representative Fukushima Daiichi Unit I boundary conditions. In particular, this study confirms that the eutectics model generally exhibits accelerated degradation and failure of fuel components, the core plate, and the lower head. The eutectics model also has a tendency to exhibit a greater degree of core degradation, greater debris mass formation, and larger debris mass ejection. Finally, the eutectics model exhibits higher maximum temperatures for fuel, cladding, particulate debris, oxidic molten pool, and metallic molten pool components than the interactive materials model; interactive materials model simulations exhibit a soft “limitation” on maximum temperatures that is related to the temperature at which material relocation occurs.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Sensitivity of UO 2 fuel performance to microstructural evolutions driven by dilute additives

Use of dilute additives to nuclear fuel is being considered to increase the security of commercial fuel management through traceability of fabricated fuel elements. Taggants, as additives are denoted when included for traceability purposes, may also improve fuel performance, as demonstrated in Cr-containing uranium dioxide as described in the literature, and they may also improve fuel safety. In fact, studies have shown that some additives affect fuel material properties such as grain size and density after sintering. Given the possible range of elements that could be used as additives, the impact of such fuel property variations on the fuel’s thermomechanical behavior becomes relevant. These effects can be evaluated through a sensitivity study of standard fuel models to analyze changes in these properties using a fuel performance code. In this work, the BISON code is being used to investigate these effects through a 2D axisymmetric model of smeared UO 2 fuel pellets and ZIRLO® cladding under realistic pressurized water reactor core irradiation conditions. Here, randomly sampled densities and grain sizes within specified ranges are used as input parameters in the simulations, and several fuel model-related outputs are evaluated. The thermomechanical response of the cladding is also addressed in this study. The simultaneous variation of both input parameters offers a more comprehensive path to identify key sensitivities. Outputs explored include temperature, fission gas release, creep, and radial stress. Results show that although most of these outputs are sensitive to grain size to a certain extent, density mainly affects fuel temperature and elastic strain. Furthermore, sensitivities can vary depending on the radial position within the fuel pellet.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Sensitivity and uncertainty of the IFR-1 BISON benchmark

The fuel performance code BISON is being used to evaluate metallic fuel for a new fast-spectrum test reactor called the Versatile Test Reactor, which is being considered for adoption by the US Department of Energy. To quantify the accuracy of BISON predictions, researchers at Oak Ridge National Laboratory have been developing a series of benchmarks based on legacy metallic fuel experiments. As part of this effort, the sensitivity of BISON predictions to variations in model inputs and the uncertainties associated with BISON predictions must be established. This paper summarizes efforts to perform a comprehensive sensitivity analysis (SA) and uncertainty quantification (UQ) on a benchmark based on the IFR-1 experiment.For the SA, at least one input was chosen from every BISON model and physics module used in the benchmark. The inputs were varied individually in a series of BISON simulations. Here, the resulting variations in benchmark predictions were normalized to calculate sensitivities. These sensitivities were then used to inform input selections for the UQ.The UQ was performed using the Monte Carlo UQ method. A literature review was conducted to estimate uncertainty distributions for the selected inputs, and values were sampled randomly from each distribution in a series of BISON simulations. Variations in the benchmark predictions were used to estimate uncertainty distributions and confidence intervals. It was found that nearly 100% of the benchmark predictions matched the corresponding legacy values within the confidence intervals. However, this is at least partially because the confidence intervals associated with benchmark predictions were wide. The uncertainty contributions of assumptions in the benchmark, experimental uncertainties, and BISON models were quantified. Some analysis was performed to identify inputs that contributed to the uncertainties. Finally, recommendations are made for future benchmark and future BISON development.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Exploring the potential of using L-Band InSAR for the mapping of flooded vegetation in tropical wetlands

Wetlands play a critical role in global water and carbon cycles, yet monitoring their water extent remains difficult, particularly beneath dense vegetation. SAR-based techniques such as backscatter thresholding are limited by complex scattering mechanisms, while fully polarimetric SAR (PolSAR) data capable of detecting doublebounce scattering remain scarce. To address these challenges, this study evaluates the potential of Interferometric SAR (InSAR) for mapping water surfaces beneath vegetation, termed flooded vegetation, using the Atrato floodplain in Colombia as a case study. We develop an automated workflow combining InSAR fringe detection with local phase homogeneity analysis and random sampling of processing parameters to generate probabilistic flooded vegetation maps. Applied to ALOS PALSAR-1 L-band image pairs from 2007–2011, the workflow captures seasonal fluctuations in flooded extent ranging from 500 to 1,500 km2. Compared to other L-band SAR inundation products, the InSAR-based maps identify broader flooded areas, with ~70% agreement in pairwise comparisons. Around 84% of detections align with existing wetland inventories and seasonal changes correspond with regional hydrological indicators, including terrestrial water storage anomalies and water gauge measurements. PolSAR analysis shows that InSAR complements backscatter-based methods by detecting inundation in areas with weak double-bounce signals. These findings suggest that combining InSAR with backscatter-based methods can improve detection of flooded vegetation, which is especially relevant for the upcoming NISAR mission that will offer frequent global L-band observations.

Coastal inundation↗

Co-Location of Cellulosic Bioethanol and Alcohol-to-Jet (ATJ) Production Facilities for Targeted Scale-Up of Sustainable Aviation Fuel (SAF) Production

Achieving aerospace industry net-zero emissions by 2050 requires rapid scaling of sustainable aviation fuel (SAF) production. Leveraging existing infrastructure, proven technologies like Alcohol-to-Jet (ATJ), and low carbon intensity (CI) feedstocks (e.g., switchgrass and miscanthus) can support this transition and help achieve near-term emissions reduction targets. This study evaluates the implications of lignocellulosic ethanol biorefinery siting and integration with petroleum refineries to produce SAF across 1000 sites randomly sampled from areas suitable for perennial grasses in the U.S. rainfed region. To better understand the logistics of material transport and handoffs, we integrated models of biomass harvest, transport, ethanol, and ATJ production in a stochastic framework based on Monte Carlo simulations to characterize SAF minimum selling price (MSP) and carbon intensity (CI), considering site-specific parameters (e.g., feedstock production, transportation, taxes, incentives). The results indicate trade-offs between MSP and CI across locations, with median MSP ranging from 7.9 to 12.8 USD·gal −1 and CI from −9.7 to 39.4 gCO 2 e·MJ −1 . Despite high estimated decarbonization costs (580 USD·tonCO 2 e −1 ), our results indicate that site-specific deployment of ATJ with low-CI feedstocks can improve sustainability outcomes. The framework provides a systematic approach to assess cost and sustainability trade-offs across locations, considering the end-to-end supply chain and supporting an informed investment in SAF production.

09 BIOMASS FUELS↗

The Lack of QBO–MJO Connection in CMIP6 Models

Observational analysis has indicated a strong connection between the stratospheric quasi–biennial oscillation (QBO) and tropospheric Madden–Julian oscillation (MJO), with MJO activity being stronger during the easterly phase than the westerly phase of the QBO. We assess the representation of this QBO–MJO connection in 30 models participating in the Coupled Model Intercomparison Project 6. While some models reasonably simulate the QBO during boreal winter, none of them capture a difference in MJO activity between easterly and westerly QBO that is larger than that which would be expected from the random sampling of internal variability. The weak signal of the simulated QBO–MJO connection may be due to the weaker amplitude of the QBO than observed, especially between 100 to 50 hPa. This weaker amplitude in the models is seen both in the QBO–related zonal wind and temperature, the latter of which is thought to be critical for destabilizing tropical convection.

54 ENVIRONMENTAL SCIENCES↗

A Machine‐Learning‐Assisted Stochastic Cloud Population Model as a Parameterization of Cumulus Convection

Abstract A machine‐learning‐assisted stochastic cloud population model is coupled with the Advanced Research Weather Research and Forecasting (WRF) model to represent fluctuations in the cloud‐base mass flux associated with the life cycles and interactions among cumulus convection cells. In this cloud population model, the size distribution and the associated cloud‐base mass flux of the convective cells are related to their previous state and to the change in the total convective area via a transition function. The convective area tendency in turn is assumed to depend on the cloud‐base mass flux that is resolved by the host WRF model. The transition function is represented by a single hidden‐layer neural network trained by the evolution of convective cell size distributions in a 1‐km grid‐spacing WRF simulation run over the Australian Monsoon region. At every grid point of the host model, the cloud population model predicts the cell size and cloud‐base mass flux distributions from which a random sample of cells is fed to an entraining parcel model that calculates precipitation as well as the associated liquid water potential temperature and total moisture tendencies. These tendencies are averaged over the cells and provided to the host model. Several regional simulations are performed over tropical and midlatitude domains to test this as a potential approach to scale‐aware parameterization. It is shown that such an approach could be a new promising path to simulating realistic precipitation statistics and propagation of precipitation associated with the Madden‐Julian Oscillation while maintaining realistic depictions of the diurnal cycle over both land and ocean.

54 ENVIRONMENTAL SCIENCES↗

Denoising Autoencoder for Reconstructing Sensor Observation Data and Predicting Evapotranspiration: Noisy and Missing Values Repair and Uncertainty Quantification

Abstract Machine learning (ML) methods applied in scientific research often deal with interrelated features in high‐dimensional data. Reducing data noise and redundancy is needed to increase prediction accuracy and efficiency especially when dealing with data from field sensors. We explored an unsupervised learning method, the denoising autoencoder (DAE), to extract the underlying data structure from noisy raw data in the context of predicting hydrologic quantities from multiple field sensors. These sensors have intrinsic instrumental noise and occasional malfunctions that cause missing values. Our DAE neural network reconstructed meteorological sensor data containing noise and missing values to predict evapotranspiration in a mountainous watershed. The DAE reconstructed the sensor variables with a mean coefficient of determination value of 0.77 across 15 dimensions representing individual sensors. It reduced variance and bias uncertainties compared to a classical autoencoder model. The reconstruction quality varied across dimensions depending on their cross‐correlation and alignment with the underlying data structure. Uncertainties arising from the model structure were overall higher than those resulting from data corruption. We attached the DAE structure to a downstream ET‐prediction neural network in three formats and achieved reasonably accurate ET predictions . The use of the DAE notably reduced variance uncertainty in ET prediction. However, excessive variance reduction may be accompanied by an increase in bias due to the intrinsic bias‐variance tradeoff. Our method of evaluating and reducing uncertainties in aggregated data from different sources can be used to improve predictive models, process understanding, and uncertainty quantification for better water resource management. Plain Language Summary We present a machine learning method, namely the denoising autoencoder, which reduces the effects of data noise and missing values typically present in scientific data sets collected through sensor measurements. This method selects the most relevant information from noisy raw data collected by the instruments and fills in missing values. To demonstrate the effectiveness of our method, we applied it to predict evapotranspiration, a hydrologic variable that represents the water moved from the land surface to the atmosphere through a combination of evaporation and plant water use (transpiration). We also used a random sampling technique (the Monte Carlo method) to compare the uncertainty in the predictions when using the raw and noisy data versus the reconstructed data. The denoising process produced more accurate predictions of evapotranspiration with less uncertainty. Improved predictions of evapotranspiration can lead to a better understanding and accounting of water budgets. This ML approach is broadly suitable for a wide variety of applications that involve noisy sensor data with missing values. Key Points We used a denoising autoencoder (DAE) neural network to reduce noise in meteorological and soil sensor observations by on average We used Monte Carlo sampling to estimate the bias and variance of all model outputs, including uncertainty sources from data and the model We attached the DAE component to a downstream neural network to predict ET with the variance reduced by , compared to that without the DAE

denoising autoencoder↗

From Points to Planes: A Workflow for Converting Three‐Dimensional Point Cloud Data Into Discrete Fracture Network Flow and Transport Models

We present the Point cLoud Algorithm for NEtwork Extraction of Discrete Fracture Networks (PLANE-DFN), a point cloud–based algorithm for automatic fracture network extraction designed to support discrete fracture network (DFN) modeling workflows. PLANE-DFN segments three-dimensional fracture planes from raw point cloud data using RANdom SAmple Consensus coupled with statistical outlier removal and density-based clustering to isolate individual fracture features. Each candidate plane is constrained against site-specific structural constraints based on strike and dip. After segmentation, each fracture is converted into a 2-D convex polygon suitable for meshing and simulation. The PLANE-DFN algorithm is validated by comparing geometric and flow and transport data against data from dfnWorks simulations with ensembles of plane-fit networks. We find that the flow and transport in plane-fit networks are comparable to dfnWorks-generated networks when realistic network geometry is maintained. The PLANE-DFN algorithm provides an automated and streamlined workflow to transform point clouds of data into DFN network geometry.

54 ENVIRONMENTAL SCIENCES↗

Diversity of visual inputs to Kenyon cells of the Drosophila mushroom body

The arthropod mushroom body is well-studied as an expansion layer representing olfactory stimuli and linking them to contingent events. However, 8% of mushroom body Kenyon cells in Drosophila melanogaster receive predominantly visual input, and their function remains unclear. Here, we identify inputs to visual Kenyon cells using the FlyWire adult whole-brain connectome. Input repertoires are similar across hemispheres and connectomes with certain inputs highly overrepresented. Many visual neurons presynaptic to Kenyon cells have large receptive fields, while interneuron inputs receive spatially restricted signals that may be tuned to specific visual features. Individual visual Kenyon cells randomly sample sparse inputs from combinations of visual channels, including multiple optic lobe neuropils. These connectivity patterns suggest that visual coding in the mushroom body, like olfactory coding, is sparse, distributed, and combinatorial. However, the specific input repertoire to the smaller population of visual Kenyon cells suggests a constrained encoding of visual stimuli.

59 BASIC BIOLOGICAL SCIENCES↗

Autonomous reinforcement learning agent for stretchable kirigami design of 2D materials

Abstract Mechanical behavior of 2D materials such as MoS 2 can be tuned by the ancient art of kirigami. Experiments and atomistic simulations show that 2D materials can be stretched more than 50% by strategic insertion of cuts. However, designing kirigami structures with desired mechanical properties is highly sensitive to the pattern and location of kirigami cuts. We use reinforcement learning (RL) to generate a wide range of highly stretchable MoS 2 kirigami structures. The RL agent is trained by a small fraction (1.45%) of molecular dynamics simulation data, randomly sampled from a search space of over 4 million candidates for MoS 2 kirigami structures with 6 cuts. After training, the RL agent not only proposes 6-cut kirigami structures that have stretchability above 45%, but also gains mechanistic insight to propose highly stretchable (above 40%) kirigami structures consisting of 8 and 10 cuts from a search space of billion candidates as zero-shot predictions.

36 MATERIALS SCIENCE↗

Automated Construction of a Photocatalysis Dataset for Water-Splitting Applications

We present an automatically generated dataset of 15,755 records that were extracted from 47,357 papers. These records contain water-splitting activity in the presence of certain photocatalysts, along with additional information about the chemical reaction conditions under which this activity was recorded. These conditions include any co-catalysts and additives that were present during water splitting, the length of time for which the photocatalytic experiment was conducted, and the type of light source used, including its wavelength. Despite the text extraction of such a wide range of chemical reaction attributes, the dataset afforded good precision (71.2%) and recall (36.3%). These figures-of-merit were calculated based on a random sample of open-access papers from the corpus. Mining such a complex set of attributes required the development of novel techniques in knowledge extraction and interdependency resolution, leveraging inter- and intra-sentence relations, which are also described in this paper. We present a new version (version 2.2) of the chemistry-aware text-mining toolkit ChemDataExtractor, in which these new techniques are included.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗