Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data scarce”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

ML-based Micro-CT SOFC Microstructure Models (from Kent 2026 Microstructural Augmentation paper)

Overview -------------------------- This repository contains datasets from the manuscript **"Enhanced Generalizability to Deep-Learning Quantification of 3D Microstructural Characteristics through Microstructurally Aware Augmentation of Scarce Data"** (*William F. Kent, Rochan Bajpai, Rachel C. Kurchin, William K. Epting, Harry W. Abernathy, Paul A. Salvador. Submitted 2026*). The methods are also described in the dissertation **Data Intensive Analysis of Solid Oxide Cell Microstructures** (*Doctoral dissertation, Carnegie Mellon University, 2025*). The datasets here are trained convolutional neural network (CNN) models for predicting key microstructural properties of solid oxide cell (SOC) electrodes from low-res, 2-channel 3D images, as well as some helpful code. The parameters for input images are provided in the paper. Sample data is provided in the file `Combined_anode_aug_dual_1k_examples` - that particular data was used to train `anode_all_aug.pth` and will work most accurately with that model. Please familiarize yourself with all caveats on accuracy and applicability, as detailed in the associated paper. Usage -------------------------- The basic usage is as follows, assuming `model_fn` is the path to the .pth file, and `X` is 2-channel input image(s) of the proper dimensions (either one image of shape `[2,12,24,24]`, or a batch of N input images of shape `[N,2,12,24,24]`): from CNN_inferencer import load_model_for_inference model = load_model_for_inference(model_fn) y_predicted = model(X) The model object automatically handles input scaling and output de-scaling based on the way the models were trained - in other words, pass in a 2-channel micro-CT image, and it will output microstructural property values in real units. ## Other model object attributes Note that model has useful attributes other than its forward pass model(X). * `model.output_descaler` - returns the output descaler object. Model does the de-scaling when generating inferences, but you may want to re-use this de-scaler on other values to e.g. compare predictions to ground truth from already-scaled training data. * `model.prop_names` - Gives the property names of the predicted y values, in order. Only exists if there's an output scaler as part of the model object, which there will be in the models provided here. ## Usage with sample data Here is a short script to use with the included sample data. from CNN_inferencer import display_predictions, load_model_for_inference, calculate_mape, parity_plot import h5py import numpy as np model_fn = 'anode_all_aug.pth' data_fn = 'Combined_anode_aug_dual_1k_examples.h5' N_samples = 200 figure_outdir = '.' model = load_model_for_inference(model_fn) with h5py.File(data_fn,'r') as f: XX = f['X'] #These are the 2-channel 3D images yy = f['y'] #These are the ground-truth microstructural properties, but they have been scaled for training - need to de-scale below N = XX.shape[0] #How many images total in the input data file #Run inferences on N_samples random samples from XX. #Run in a batch, much more efficient than one at a time. ii = np.random.choice(N,N_samples,replace=False) ii.sort() y_pred = model(XX[ii]) #Get the original/true (but normalized/scaled) values from the training dataset... #Because they were normalized, they are not in real units yet. So let's also de-scale them using model.output_scaler. y_true = model.output_scaler.transform(yy[ii]) #Let's display actual values for just 5 random ones for i in np.random.choice(N_samples,5,replace=False): display_predictions(y_true[i], y_pred[i], model.prop_names) #Make parity plots for each property (ground truth vs predicted values) #Also label each plot with the mean abs. percent error (MAPE) of the predicted values for i,key in enumerate(model.prop_names): mape = calculate_mape(y_true[:,i], y_pred[:,i]) parity_plot(y_true[:,i], y_pred[:,i], figure_outdir, key, extra_title=f' ({mape:.2f}% MAPE)')

3D microstructure↗

Satellite Embedding-Based Population Imputation for Areas with Missing Building Footprint Data: A Computer Vision-Based Approach

High-resolution population modeling is important for supporting effective decision-making across diverse sectors. LandScan Mosaic generates population estimates at the level of individual buildings and aggregates them to 3 arc-second grids, and this approach performs well in regions where building footprint data are comprehensive and reliable. However, large portions of the globe still suffer from incomplete, sparse, or entirely missing building stock datasets, creating a structural limitation for strictly building-based population models. To address this research gap, this study proposes a computer vision-based framework that employs Google Earth Engine satellite embeddings and UNet, which allows us to directly impute grid-level population estimates in building-data-deficient areas. Applied to Taiwan as a case study, the framework achieved strong predictive performance with R$^{2}$ of 0.89, RMSE of 18.70, and MAE of 8.41, outperforming traditional machine learning approaches. Notably, the proposed framework effectively addressed building false-positive errors inherent in Global Human Settlement Layer (GHSL) data, correctly identifying uninhabited areas that were erroneously classified as populated. The framework also offers significant advantages for global population mapping, particularly in terms of scalability and temporal consistency, thereby extending the coverage and accuracy of high-resolution population products in data-scarce regions worldwide. Urban planners, decision makers, and related stakeholders can obtain granular population distributions to support more accurate and targeted infrastructure investment, service delivery, resource allocation, and risk assessment decisions.

97 MATHEMATICS AND COMPUTING↗

Evaluation of interactive and prescribed agricultural ammonia emissions for simulating atmospheric composition in CAM-chem

Abstract. Ammonia (NH3) plays a central role in the chemistry of inorganic secondary aerosols in the atmosphere. The largest emission sector for NH3 is agriculture, where NH3 is volatilized from livestock wastes and fertilized soils. Although the NH3 volatilization from soils is driven by the soil temperature and moisture, many atmospheric chemistry models prescribe the emission using yearly emission inventories and climatological seasonal variations. Here we evaluate an alternative approach where the NH3 emissions from agriculture are simulated interactively using the process model FANv2 (Flow of Agricultural Nitrogen, version 2) coupled to the Community Atmospheric Model with Chemistry (CAM-chem). We run a set of 6-year global simulations using the NH3 emission from FANv2 and three global emission inventories (EDGAR, CEDS and HTAP) and evaluate the model performance using a global set of multi-component (atmospheric NH3 and NH4+, and NH4+ wet deposition) in situ observations. Over East Asia, Europe and North America, the simulations with different emissions perform similarly when compared with the observed geographical patterns. The seasonal distributions of NH3 emissions differ between the inventories, and the comparison to observations suggests that both FANv2 and the inventories would benefit from more realistic timing of fertilizer applications. The largest differences between the simulations occur over data-scarce regions. In Africa, the emissions simulated by FANv2 are 200 %–300 % higher than in the inventories, and the available in situ observations from western and central Africa, as well as NH3 retrievals from the Infrared Atmospheric Sounding Interferometer (IASI) instrument, are consistent with the higher NH3 emissions as simulated by FANv2. Overall, in simulating ammonia and ammonium concentrations over regions with detailed regional emission inventories, the inventories based on these details (HTAP, CEDS) capture the atmospheric concentrations and their seasonal variability the best. However these inventories cannot capture the impact of meteorological variability on the emissions, nor can these inventories couple the emissions to the biogeochemical cycles and their changes with climate drivers. Finally, we show with sensitivity experiments that the simulated time-averaged nitrate concentration in air is sensitive to the temporal resolution of the NH3 emissions. Over the CASTNET monitoring network covering the US, resolving the NH3 emissions hourly instead monthly reduced the positive model bias from approximately 80 % to 60 % of the observed yearly mean nitrate concentration. This suggests that some of the commonly reported overestimation of aerosol nitrate over the US may be related to unresolved temporal variability in the NH3 emissions.

54 ENVIRONMENTAL SCIENCES↗

The suitability of differentiable, physics-informed machine learning hydrologic models for ungauged regions and climate change impact assessment

As a genre of physics-informed machine learning, differentiable process-based hydrologic models (abbreviated as δ or delta models) with regionalized deep-network-based parameterization pipelines were recently shown to provide daily streamflow prediction performance closely approaching that of state-of-the-art long short-term memory (LSTM) deep networks. Meanwhile, δ models provide a full suite of diagnostic physical variables and guaranteed mass conservation. Here, we ran experiments to test (1) their ability to extrapolate to regions far from streamflow gauges and (2) their ability to make credible predictions of long-term (decadal-scale) change trends. We evaluated the models based on daily hydrograph metrics (Nash–Sutcliffe model efficiency coefficient, etc.) and predicted decadal streamflow trends. For prediction in ungauged basins (PUB; randomly sampled ungauged basins representing spatial interpolation), δ models either approached or surpassed the performance of LSTM in daily hydrograph metrics, depending on the meteorological forcing data used. They presented a comparable trend performance to LSTM for annual mean flow and high flow but worse trends for low flow. For prediction in ungauged regions (PUR; regional holdout test representing spatial extrapolation in a highly data-sparse scenario), δ models surpassed LSTM in daily hydrograph metrics, and their advantages in mean and high flow trends became prominent. In addition, an untrained variable, evapotranspiration, retained good seasonality even for extrapolated cases. The δ models' deep-network-based parameterization pipeline produced parameter fields that maintain remarkably stable spatial patterns even in highly data-scarce scenarios, which explains their robustness. Combined with their interpretability and ability to assimilate multi-source observations, the δ models are strong candidates for regional and global-scale hydrologic simulations and climate change impact assessment.

54 ENVIRONMENTAL SCIENCES↗

Use of Very High-Resolution Optical Data for Landslide Mapping and Susceptibility Analysis Along the Karnali Highway, Nepal

The Karnali highway is a vital transport link and the only primary roadway that connects the remote Karnali region to the lowlands in Mid-Western Nepal. Every year there are reports of landslides blocking the road, making this area largely inaccessible. However, little effort has focused on systematically identifying landslides and landslide-prone areas along this highway. In this study, landslides were mapped with an object-based approach from very high-resolution optical satellite imagery obtained by the DigitalGlobe constellation in 2012 and PlanetScope in 2018. Landslides ranging from 10 to 30,496 sq.m were detected within a 3 km buffer along the highway. Most of the landslides were located at lower elevations (between 500–1500 m) and on steep south-facing slopes. Landslides tended to cluster closer to the highway, near drainage channels and away from faults. Landslides were also most prevalent within the Kuncha Formation geologic class, and the forested and agricultural land cover classes. A susceptibility map was then created using a logistic regression methodology to highlight patterns in landslide activity. The landslide susceptibility map showed a good prediction rate with an area under the curve (AUC) of 0.90. A total of 33% of the study arealies in high/very high susceptibility zones. The map highlighted the lower elevated areas between Bangesimal and Manma towns with the Kuncha Formation geologic class as being the most hazardous. The banks of the Karnali River, its tributaries and areas near the highway were also highly susceptible to landslides. The results highlight the potential of very high-resolution optical imagery for documenting detailed spatial information on landslide occurrence, which enables susceptibility assessment in remote and data scarce regions such as the Karnali highway.

Pukar Amatya↗

Tonlé Sap Food Security and Agriculture: Evaluating the Effects of Land Use and Hydrological Change on Ecosystem Vitality using Remotely-Sensed Data in the Tonlé Sap Lake Basin

Tonlé Sap Lake, the largest lake in Southeast Asia, is a critical source of fish and freshwater resources for the region. The health of this freshwater system is under pressure from accelerating dam construction, intensifying agriculture, deforestation, and changing climate patterns, forcing tradeoffs between immediate food security and the long-term vitality and productivity of the ecosystem. Efficient freshwater system monitoring is crucial to navigating these challenges. In collaboration with Conservation International, the Cambodian Ministry of Water Resources and Meteorology, and the Tonlé Sap Authority, we developed and tested remotely-sensed proxies for sub-indicators of the Freshwater Health Index (FHI), which is typically calculated using in situ datasets. We used landcover datasets derived from Landsat 5 Thematic Mapper (TM), Landsat 7 Enhanced Thematic Mapper Plus (ETM+), Landsat 8 Operational Land Imager (OLI), PROBA-V Vegetation sensor (VGT), Sentinel-2 Multispectral Imager (MSI), Advanced Very High Resolution Radiometer (AVHRR), and Envisat Medium Resolution Imaging Spectrometer (MERIS) as inputs to calculate land cover naturalness and bank modification. Additionally, we created a lake-level time series using a collection of altimetry data sources to estimate deviation from natural flow. We observed a decrease in landcover naturalness and a breakdown in the volume and regularity of annual lake levels from 2000-2020, reflecting increased pressure on water supply and agricultural productivity. At least 8% of forested areas in the basin were lost and rice harvest intensity increased over the course of the study period. These results will help our partners make informed decisions regarding freshwater management. Furthermore, our remotely-sensed FHI analysis can be replicated in other regions, providing decision makers with a snapshot of freshwater health in data-scarce environments.

Marco Vallejos↗

Soil Moisture Estimation in South Asia via Assimilation of SMAP Retrievals

A soil moisture retrieval assimilation framework is implemented across South Asia in an attempt to improve regional soil moisture estimation as well as to provide a consistent regional soil moisture dataset. This study aims to improve the spatiotemporal variability of soil moisture estimates by assimilating Soil Moisture Active Passive (SMAP) near-surface soil moisture retrievals into a land surface model. The Noah-MP (v4.0.1) land surface model is run within the NASA Land Information System software framework to model regional land surface processes. NASA Modern-Era Retrospective Analysis for Research and Applications (MERRA2) and Global Precipitation Measurement (GPM) Integrated Multi-satellitE Retrievals (IMERG) provide the meteorological boundary conditions to the land surface model. Assimilation is carried out using both cumulative distribution function (CDF)-corrected (DA-CDF) and uncorrected SMAP retrievals (DA-NoCDF). CDF matching is applied to correct the statistical moments of the SMAP soil moisture retrieval relative to the land surface model. Comparison of assimilated and model-only soil moisture estimates with publicly available in situ measurements highlights the relative improvement in soil moisture estimates by assimilating SMAP retrievals. Across the Tibetan Plateau, DA-NoCDF reduced the mean bias and RMSE by 8.4 % and 9.4 %, even though assimilation only occurred during less than 10 % of the study period due to frozen (or partially frozen) soil conditions. The best goodness-of-fit statistics were achieved for the IMERG DA-NoCDF soil moisture experiment. The general lack of publicly available in situ measurements across irrigated areas limited a domain-wide direct model validation. However, comparison with regional irrigation patterns suggested correction of biases associated with an unmodeled hydrologic phenomenon (i.e., anthropogenic influence via irrigation) as a result of SMAP soil moisture retrieval assimilation. The greatest sensitivity to assimilation was observed in cropland areas. Improvements in soil moisture potentially translate into improved spatiotemporal patterns of modeled evapotranspiration, although limited influence from soil moisture assimilation was observed on modeled processes within the carbon cycle such as gross primary production. Improvement in fine-scale modeled estimates by assimilating coarse-scale retrievals highlights the potential of this approach for soil moisture estimation over data-scarce regions.

Jawairia Ahmad↗

Utilizing Earth Observations to Understand Landscape Patterns and Assist in Wildlife Management in Iona National Park, Angola

Following the end of the Angolan Civil War (1975-2002), human habitation in Iona National Park has grown exponentially, as has the livestock population. An ongoing drought beginning in 2017 has brought people, livestock, and wildlife into increasing competition for resources within the park. This study used Earth observation data, primarily Landsat and Sentinel imagery, to examine landscape trends to improve wildlife preservation approaches in Iona National Park, Angola. In collaboration with the NGO African Parks, we developed a robust land use and land cover (LULC) classification model using remote sensing data to augment sparse ground-based data in this arid land region. We used Google Earth Engine and a random forest classifier to map vegetation types, water bodies, and potential wildlife habitats. This analysis resulted in a high spatial resolution LULC time-series between 1984-2023, highlighting critical periods of socioecological change over the past 40 years. These results increased the partner’s ability to make scientifically grounded decisions about resource allocation and conservation priorities. This analysis supports the feasibility of applying remote sensing techniques coupled with machine learning models in dry regions, where standard survey methods are frequently limited by accessibility and resource availability. However, we identified limitations in ground-truth data and the difficulty of recognizing certain vegetation types in arid areas. Despite these limitations, the study demonstrated Earth observations' ability to transform wildlife management techniques in distant and data-scarce locations, providing a reproducible foundation for similar ecosystems around the world.

Emmanuel Aklie↗

Gaussian processes meet NeuralODEs: a Bayesian framework for learning the dynamics of partially observed systems from scarce and noisy data

We present a machine learning framework (GP-NODE) for Bayesian model discovery from partial, noisy and irregular observations of nonlinear dynamical systems. The proposed method takes advantage of differentiable programming to propagate gradient information through ordinary differential equation solvers and perform Bayesian inference with respect to unknown model parameters using Hamiltonian Monte Carlo sampling and Gaussian Process priors over the observed system states. This allows us to exploit temporal correlations in the observed data, and efficiently infer posterior distributions over plausible models with quantified uncertainty. The use of the Finnish Horseshoe as a sparsity-promoting prior for free model parameters also enables the discovery of parsimonious representations for the latent dynamics. A series of numerical studies is presented to demonstrate the effectiveness of the proposed GP-NODE method including predator–prey systems, systems biology and a 50-dimensional human motion dynamical system. This article is part of the theme issue ‘Data-driven prediction in dynamical systems’.

Science & Technology - Other Topics↗

Emulators for Scarce and Noisy Data: Application to Auxiliary-Field Diffusion Monte Carlo for Neutron Matter

Understanding the equation of state (EOS) of pure neutron matter is necessary for interpreting multimessenger observations of neutron stars. Reliable data analyses of these observations require well-quantified uncertainties for the EOS input, ideally propagating uncertainties from nuclear interactions directly to the EOS. This, however, requires calculations of the EOS for a prohibitively larger number of nuclear Hamiltonians, solving the nuclear many-body problem for each one. Quantum Monte Carlo methods, such as auxiliary-field diffusion Monte Carlo (AFDMC), provide precise and accurate results for the neutron matter EOS, but they are very computationally expensive, making them unsuitable for the fast evaluations necessary for uncertainty propagation. Here, we employ parametric matrix models to develop fast emulators for AFDMC calculations of neutron matter and use them to directly propagate uncertainties of coupling constants in the Hamiltonian to the EOS. As these uncertainties include estimates of the effective field theory truncation uncertainty, this approach provides robust uncertainty estimates for use in astrophysical data analyses. In conclusion, this Letter will enable novel applications such as using astrophysical observations to put constraints on coupling constants for nuclear interactions.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Use of Machine Learning on PMU Data for Transmission System Fault Analysis

Synchrophasor technology has been used for monitoring, control, and protection of bulk power system for over 10 years. Deployment of phasor measurement units (PMUs) in the USA power system has surpassed 3000 units installed in the transmission substations as stand-alone intelligent electronic devices (IEDs) or as a software add-on to other devices such as digital protective relays (DPRs) or digital fault recorders (DFRs). By now, thousands of terabytes of PMU data may have been captured and stored by various transmission system operators (TSOs) and independent system operators (ISOs). This creates an opportunity to deploy advanced machine learning (ML) techniques to detect and classify faults recorded by PMUs automatically to be used by the system operators for rapid, critical decision-making when manual analysis of the past or unfolding events is not feasible. In this paper we offer a brief background on how the automated fault analysis may be done using DPR and/or DFR data, and compare some of the legacy approaches to the new ML approaches in the context of the system-wide PMU recordings. We then offer insights from developing practical ML solutions that have been applied on field recordings captured by close to 450 PMUs from all three US interconnections (Western, Eastern and ERCOT) over two years (2016-2017). We identify and illustrate ML challenges we addressed: inaccurate data, data with scarce and temporally imprecise fault labels, data recorded by PMUs sparsely located at substations resulting in the fault records taken afar from the ends of the faulted lines, data containing only positive sequence values, and data taken at different voltage levels. We then illustrate the ML model results for fault analysis under different application scenarios. The novelty of this study is not only in the design, implementation, and performance analysis of the ML algorithms, but also in the use of advanced fault modelling and simulation approaches to improve the training results when developing supervised ML models for fault detection and classification. Extensive simulations of faults were conducted on a 14-bus power system to create a training dataset with over 1400 accurately labelled faults. This dataset was applied to enhance the accuracy of fault detection and classification of machine learning-based models trained with small number of labelled faults in large datasets recorded in the grid interconnections ranging from 5,000 to 70,000 buses.

Synchrophasors, Machine Learning, Fault Analysis, ↗

Operator inference with roll outs for learning reduced models from scarce and low-quality data

Data-driven modeling has become a key building block in computational science and engineering. However, data that are available in science and engineering are typically scarce, often polluted with noise and affected by measurement errors and other perturbations, which makes learning the dynamics of systems challenging. Here, in this work, we propose to combine data-driven modeling via operator inference with the dynamic training via roll outs of neural ordinary differential equations. Operator inference with roll outs inherits interpretability, scalability, and structure preservation of traditional operator inference while leveraging the dynamic training via roll outs over multiple time steps to increase stability and robustness for learning from low-quality and noisy data. Numerical experiments with data describing shallow water waves and surface quasi-geostrophic dynamics demonstrate that operator inference with roll outs provides predictive models from training trajectories even if data are sampled sparsely in time and polluted with noise of up to 10%.

97 MATHEMATICS AND COMPUTING↗

Aerodynamic Data Fusion Toward the Digital Twin Paradigm

This paper considers the fusion of two aerodynamic data sets originating from differing types of physical or computer experiments. This paper specifically addresses the fusion of 1) noisy and in-complete fields from wind-tunnel measurements and 2) deterministic but biased fields from numerical simulations. These two data sources are fused in order to estimate the true field that best matches measured quantities that serve as the ground truth. For example, two sources of pressure fields about an aircraft are fused based on measured forces and moments from a wind-tunnel experiment. A fundamental challenge in this problem is that the true field is unknown and cannot be estimated with 100% certainty. A Bayesian framework is employed to infer the true fields conditioned on measured quantities of interest; essentially a statistical correction to the data is performed. The fused data may then be used to construct more accurate surrogate models suitable for early stages of aerospace design. An extension of the proper orthogonal decomposition with constraints is also introduced to solve the same problem. In this work, both methods are demonstrated on fusing the pressure distributions for flow past the RAE2822 airfoil and the Common Research Model wing at transonic conditions. Comparison of both methods reveals that the Bayesian method is more robust when data are scarce and capable of also accounting for uncertainties in the data. Furthermore, given adequate data, the proper-orthogonal-decomposition-based and Bayesian approaches lead to surprisingly similar results.

42 ENGINEERING↗

Report to NCSP on 2008 DANCE measurements of 233 U($\eta$,$\gamma$)

Uranium-233 plays an important role in the Th-U fuel cycle, with substantial production in the cycle. This cycle has been proposed as an alternative to the U-Pu fuel cycle due to its reduced production of transuranic elements. An accurate measurement of the 233 U($\eta$,$\gamma$) cross section is required by the National Criticality Safety Program (NCSP) to complete the neutron-induced cross section data, where experimental capture cross section data are scarce and were measured decades ago. The most recent capture cross section data available in the literature were measured in 2007 at the n_TOF facility (CERN); in the 60s measurements were performed at Rensselaer Polytechnic Institute (RPI) and at LANL. Finally, as reported by ORNL, a new evaluation with a revised (renormalized) fission cross section is needed on 233 U. The challenge for this measurement lies in the difficulty of measuring the capture cross section data in the competing fission background, as the fission cross section is around one order of magnitude larger than the capture cross section for 233 U. The accuracy of a capture cross section measurement depends on discrimination between $\gamma$’s produced in capture and fission reactions, for which an experimental setup combining capture and fission detectors is needed. For the ($\eta$,$\gamma$) cross section measurement at LANSCE, this discrimination is achieved by combining the Detector for Advanced Neutron Capture Experiments (DANCE), to measure $\gamma$’s from capture reactions, with a Parallel Plate Avalanche Counter (PPAC) to tag the $\gamma$’s produced by fission. This method was successfully used to measure 235 U and 239 Pu capture cross sections. In these measurements, the neutron capture cross section was determined in a large fission background well above 100 keV. As part of the NCSP nuclear data effort, we have looked at past DANCE measurements on 233 U($\eta$,$\gamma$) and evaluated whether existing data is adequate to apply this technique.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Describing Point Defect Topology in 2D Energy Materials Through Computer Vision

Point defects such as vacancies and impurity atoms strongly impact the performance of 2D materials. Traditional efforts often rely on manual detection, a process that is time-intensive, prone to human error, and challenging to scale. Here we leverage machine learning (ML) methods to identify and quantify vacancies within 2D transition metal carbides (Ti3C2, MXenes), aiming to expedite detection while improving accuracy. MXenes exhibit valuable defect-defined electrochemical properties, but we currently lack statistical understanding of defect topology needed to fully harness these materials. Here we employ a convolutional neural network for semantic segmentation of experimental MXene images, opening an opportunity to conduct a rigorous statistical study on defect hierarchy while investigating local relaxation in the lattice. We show how the integration of ML can yield fundamental insight into point defects, providing a powerful tool that will play an increasingly crucial role in the future of materials science. ML is often not just a matter of straightforward application, and pretrained models proved ineffective in this case. Instead, we trained our own neural network (NN) and applied data augmentation techniques and fine-tuning to the training dataset. Since labeled microscopy data is often scarce, we developed training data from a previously published wide-frame MXene image, using customized Gaussian fitting to locate atomic positions. Our trained model was then applied to a large dataset of experimental images, enabling a statistical study of defect configurations across three samples prepared with different HF etchant concentrations (5%, 9.1%, and 12.5%), as shown in Fig. 1. This also allowed us to investigate local strain around vacancies, though we find that we are limited by the precision of measurements using high-angle annular dark field (HAADF) images, as shown in Fig. 2. This study demonstrates how ML enables large-scale, quantitative analysis of atomic defects - an otherwise infeasible task with traditional methods. While our NN was specialized for Ti3C2 MXenes, the pipeline we developed provides a foundation for future ML models tailored to other materials. Ultimately, we envision embedding the NN onto the microscope to give real-time feedback to the user. To make this a reality, continued work is necessary to fully understand the NN's capabilities and limitations. This study gets one step closer to our goals of automated experimentation moving away from traditional methods of manual labeling. As ML capabilities advance, we hope to continue adapting and applying these techniques in microscopy.

2D materials↗

Investigation of Benchmark $k$ eff Sensitivity and Uncertainty for 239 Pu fission in Specific Energy Ranges

Nuclear data at intermediate energies (from 1 to 100s of keV) are evaluated based on scarce differential data and theory unable to capture physics’ expected structure. There is also a lack of integral data. This is a known deficiency and is challenging to address. Calculated effective multiplication factor, k eff , values for intermediate energy experiments are ~25× further from experiment than for fast energies and are often well outside the experimental uncertainties. The goal of the PARADIGM (PARallel Approach of Differential and InteGral Measurements) project is to significantly re duce the uncertainties of intermediate energy nuclear data for 239 Pu. To this end, PARADIGM simultaneously optimizes experiments at both the Los Alamos Neutron Science Center (LANSCE) and National Criticality Experiments Research Center (NCERC). The combined set of data will inform new intermediate-energy nuclear data. By execution of differential and integral experiments, establishment of new theory, and undertaking nuclear data evaluation in parallel, the timeline to deliver improved nuclear data to users will be reduced significantly that is to three years. For the PARADIGM project, it was decided to optimize an integral experiment for two neutron energy ranges, within the full intermediate energy range. The low energy range goes from 1 to 30 keV, while the higher energy range goes from 30 to 600 keV. This work focuses on nuclear data sensitivities and uncertainties for 239 Pu fission for existing experiments in the International Criticality Safety Benchmark Evaluation Project (ICSBEP). When designing new experiments, it is important to understand what benchmarks currently exist. For a more traditional experiment design (in which a specific application model(s) exists), comparisons would be made between the application model(s) and existing benchmarks. For PARADIGM, there is no specific application model, but instead the specific nuclear data reaction and energy ranges of interest can be explored for existing benchmarks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Understanding collective human movement dynamics during large-scale events using big geosocial data analytics

Conventional approaches for modeling human mobility pattern often focus on human activity and movement dynamics in their regular daily lives and cannot capture changes in human movement dynamics in response to large-scale events. With the rapid advancement of information and communication technologies, many researchers have adopted alternative data sources (e.g., cell phone records, GPS trajectory data) from private data vendors to study human movement dynamics in response to large-scale natural or societal events. Big geosocial data such as georeferenced tweets are publicly available and dynamically evolving as real-world events are happening, making it more likely to capture the real-time sentiments and responses of populations. However, precisely-geolocated geosocial data is scarce and biased toward urban population centers. In this research, we developed a big geosocial data analytical framework for extracting human movement dynamics in response to large-scale events from publicly available georeferenced tweets. The framework includes a two-stage data collection module that collects data in a more targeted fashion in order to mitigate the data scarcity issue of georeferenced tweets; in addition, a variable bandwidth kernel density estimation(VB-KDE) approach was adopted to fuse georeference information at different spatial scales, further augmenting the signals of human movement dynamics contained in georeferenced tweets. To correct for the sampling bias of georeferenced tweets, we adjusted the number of tweets for different spatial units (e.g., county, state) by population. To demonstrate the performance of the proposed analytic framework, we chose an astronomical event that occurred nationwide across the United States, i.e., the 2017 Great American Eclipse, as an example event and studied the human movement dynamics in response to this event. Finally, this analytic framework can easily be applied to other types of large-scale events such as hurricanes or earthquakes.

54 ENVIRONMENTAL SCIENCES↗