Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hierarchical data change analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

43 records · Page 3

Mitigating Algorithmic Bias in Cancer Site Classification Models

Purpose Integrating artificial intelligence in cancer diagnostics has improved tumor classification beyond rule-based systems. Despite these advancements, these models may still encode demographic biases. We conducted a large-scale, applied bias-probing study of a deep learning–based cancer site classifier to quantify race information encoded in document embeddings. We then evaluated how performance changes when race-correlated embedding dimensions are removed in a post-training sensitivity analysis. Methods The cancer site classifier was trained using 3.5 million electronic cancer pathology reports from six of the National Cancer Institute's SEER registries. We trained a hierarchical self-attention network to generate 400-dimensional document embeddings. These embeddings were used to train two downstream, gradient-boosted decision tree classifiers: one to classify the cancer sites and another to predict racial categories. We identified overlapping features by intersecting the top 50 feature-importance rankings from the site and race models and computed their cumulative feature importance in each model. As a post hoc sensitivity analysis, we progressively pruned these overlapping dimensions, retrained the site model, and compared overall macro-F1 and accuracy, race-stratified macro-F1, and group fairness metrics on the basis of demographic parity and equalized odds before and after pruning. Results The analysis revealed minimal feature overlap between the cancer site and race prediction models, and the cumulative importance scores indicated a negligible influence of racial information on clinical predictions. Post-training pruning of overlapping features did not compromise the models' diagnostic accuracy, with a 0.07% loss in accuracy. Conclusion Our findings demonstrate that HiSAN-generated embeddings from SEER data can be used effectively in cancer site classification without significant demographic bias influencing the outcomes. Post-training pruning therefore functions as a practical audit and sensitivity check.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)↗

Bayesian calibration of irradiated graphite property models under high temperatures

Graphite under high temperatures and irradiation is central to advanced reactors. We develop a Bayesian calibration framework for graphite property models that explicitly represents model-data mismatch via a Gaussian-process discrepancy. The approach propagates uncertainty from parameters, experimental noise, and model form, with a hierarchical variance structure to capture group and cross-group noise. Using two predictive models across five grades (IG-110, NBG-18, PCEA, NBG-17, 2114) and four properties-irradiation-induced dimension change, creep, Young’s modulus change ratio, and coefficient of thermal expansion change ratio-we obtain average predictive-error reductions of 54%, 65%, 17%, and 17% when discrepancy is included. We illustrate engineering impact with a multiphysics model of a very-high-temperature reactor prismatic reflector brick, analyzing stresses under high fluence and temperature. Accounting for model discrepancy markedly improves predictive accuracy and provides a robust basis for reliable graphite component design in advanced reactors.

36 - MATERIALS SCIENCE↗

PS3: The Pheno-Synthesis software suite for integration and analysis of multi-scale, multi-platform phenological data

Phenology is the study of recurring plant and animal life-cycle stages which can be observed across spatial and temporal scales that span orders of magnitude (e.g., organisms to landscapes). The variety of scales at which phenological processes operate is reflected in the range of methods for collecting phenologically relevant data, and the programs focused on these collections. Consideration of the scale at which phenological observations are made, and the platform used for observation, is critical for the interpretation of phenological data and the application of these data to both research questions and land management objectives. However, there is currently little capacity to facilitate access, integration and analysis of cross-scale, multi-platform phenological data. This paper reports on a new suite of software and analysis tools – the “Pheno-Synthesis Software Suite,” or PS3 – to facilitate integration and analysis of phenological and ancillary data, enabling investigation and interpretation of phenological processes at scales ranging from organisms to landscapes and from days to decades. We use PS3 to investigate phenological processes in a semi-aride, mixed shrub-grass ecosystem, and find that the apparent importance of seasonal precipitation to vegetation activity (i.e., “greenness”) is affected by the scale and platform of observation. We end by describing potential applications of PS3 to phenological modeling and forecasting, understanding patterns and drivers of phenological activity in real-world ecosystems, and supporting agricultural and natural resource management and decision-making.

54 ENVIRONMENTAL SCIENCES↗

Local-scale heterogeneity of soil thermal dynamics and controlling factors in a discontinuous permafrost region

In permafrost regions, the strong spatial and temporal variability in soil temperature cannot be explained by the weather forcing only. Understanding the local heterogeneity of soil thermal dynamics and their controls is essential to understand how permafrost systems respond to climate change and to develop process-based models or remote sensing products for predicting soil temperature. In this study, we analyzed soil temperature dynamics and their controls in a discontinuous permafrost region on the Seward Peninsula, Alaska. We acquired one-year temperature time series at multiple depths (at 5 or 10 cm intervals up to 85 cm depth) at 45 discrete locations across a 2.3 km 2 watershed. We observed a larger spatial variability in winter temperatures than that in summer temperatures at all depths, with the former controlling most of the spatial variability in mean annual temperatures. We also observed a strong correlation between mean annual ground temperature at a depth of 85 cm and mean annual or winter season ground surface temperature across the 45 locations. We demonstrate that soils classified as cold, intermediate, or warm using hierarchical clustering of full-year temperature data closely match their co-located vegetation (graminoid tundra, dwarf shrub tundra, and tall shrub tundra, respectively). We show that the spatial heterogeneity in soil temperature is primarily driven by spatial heterogeneity in snow cover, which induces variable winter insulation and soil thermal diffusivity. These effects further extend to the subsequent summer by causing variable latent heat exchanges. Finally, we discuss the challenges of predicting soil temperatures from snow depth and vegetation height alone by considering the complexity observed in the field data and reproduced in a model sensitivity analysis.

54 ENVIRONMENTAL SCIENCES↗

High-Quality Revision of the Israeli Seismic Bulletin

Seismic bulletins, with trustworthy phase picks, origin times, and source locations are key for regional seismic studies, such as travel-time (TT) tomography, attenuation tomography, and anisotropy studies. To lay the groundwork for such studies in Israel, we revised the seismic bulletin of Israel and the surrounding area and obtained a trustworthy TT data set. From the earthquake and explosion bulletins of the Geophysical Institute of Israel, we compiled a starting data set of about 123,000 earthquakes and explosions that occurred during the past 40 yr. After screening out the poorly recorded events, we were left with a data set of ~38,000 well-recorded events. We then revised the remaining data set in two consecutive steps. In the first, we reviewed and updated station metadata, including changes in station metadata parameters over time. In the second step, we jointly relocated a list of selected seismic events, using the Bayesian hierarchical location software package (BayesLoc) of Myers et al. (2007) that performs joint relocation of multiple events. We observed striking dissimilarities between the spatial distributions of the newly relocated catalog and the initial locations. Although the depth distribution of the starting catalog is trimodal with peaks at 0, 5, and 10 km, the distribution in this study is unimodal, with a broad peak between 7.5 and 12.5 km. By differencing the observed arrival times and the origin times obtained through relocation with BayesLoc, we obtained a revised TT database that consists of 261,336 Pg, 132,876 Pn, 114,816 Sg, and 60,394 Sn arrivals, from a set of 30,458 jointly relocated seismic sources. In this work, we compared prerevision and postrevision TTs as a function of epicentral distance and concluded that the revised data set contains far fewer outliers and inconsistencies than the original data set. The revised TT data set may be used for seismic studies, such as TT tomography, attenuation tomography, and anisotropy studies.

58 GEOSCIENCES↗

Transportation Hub Infrastructure Expansion: Decision Support Under Uncertainty

The Athena project (www.athena-mobility.org) has worked to investigate the relationship between the Dallas-Fort Worth Airport (DFW) and the greater Dallas area in order to better understand and therefore better inform future decision-making regarding the critical infrastructure that influence mobility between the airport and the city. Through this work, infrastructure related to curbside pickup and drop-off, parking, public transit, and the road network congestion were identified as critical to the operation of the DFW transportation hub. The infrastructure analysis and expansion aspect of the Athena project is focused on the restructuring of the CTA curb as a hierarchical curb and the building or repurposing of parking infrastructure as the interplay between these two areas. Many sources of uncertainty exist that may impact future airport and transportation hub operations, such as passenger volume growth, population demographic changes over time, electric vehicle (EV) adoption rates, and autonomous vehicle (AV) adoption rates. Due to these sources of uncertainty, we have selected for our research a modeling framework that can capture various types of uncertainty and hedge against those uncertainties in the optimization process. We analyze road network and curb congestion, the rise of transportation networking companies, trends in parking usage, existing policies around this infrastructure, airport revenue streams, and other contributing factors to enable infrastructure decision making with less uncertainty. To accomplish this wholistic analysis, we have developed a novel multi-stage, multi-period stochastic optimization model which considers the airport's decisions from 2025-2045 under different possible future macro trajectories and day-to-day variations in operational conditions captured as "annual representation of operations" scenarios with respective probabilities. This model has also been designed to leverage the outputs of various efforts under the Athena project to create a combined decision framework for infrastructure decisions. These various efforts include the route optimization model, the ASPIRES simulation, the mode choice model, and the SUMO traffic simulation. Our computational experiments of this system at scale have resulted in a working version of our infrastructure model which enables the explicit representation and consideration of various sources of uncertainty in the decision process to enable robust, flexible decision-making. This model has been effectively run on NREL's HPC system, Eagle, with large numbers of stochastic scenarios and shows promise as a scalable tool for robust consideration of uncertainties in airport planning. We have tested our model using 30,240 operational circumstances in total, resulting in a problem with more 200 million variables. This model was solved in several different configurations, and a workflow to simulate the performance of the infrastructure model results was developed and deployed. In general, our results indicate that a combination of remote parking, remote curb infrastructure, and dynamic pricing can generate revenue, reduce emissions, accommodate emerging technologies such as AVs and EVs, and manage airport passenger growth over time. We note the success of the proposed strategy depends on the data collection and forecasting abilities of DFW. We have also seen that the AV adoption by TNCs might necessitate larger amounts of remote curb. The results of this work inform strategies for airport infrastructure decision making, as well as demonstrate the value of an adaptable model, but also indicate that there are avenues remaining where further research would be of value.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Nominal 30-M Cropland Extent Map of Continental Africa by Integrating Pixel-Based and Object-Based Algorithms Using Sentinel-2 and Landsat-8 Data on Google Earth Engine

A satellite-derived cropland extent map at high spatial resolution (30-m or better) is a must for food and water security analysis. Precise and accurate global cropland extent maps, indicating cropland and non-cropland areas, is a starting point to develop high-level products such as crop watering methods (irrigated or rainfed), cropping intensities (e.g., single, double, or continuous cropping), crop types, cropland fallows, as well as assessment of cropland productivity (productivity per unit of land), and crop water productivity (productivity per unit of water). Uncertainties associated with the cropland extent map have cascading effects on all higher-level cropland products. However, precise and accurate cropland extent maps at high spatial resolution over large areas (e.g., continents or the globe) are challenging to produce due to the small-holder dominant agricultural systems like those found in most of Africa and Asia. Cloud-based Geospatial computing platforms and multi-date, multi-sensor satellite image inventories on Google Earth Engine offer opportunities for mapping croplands with precision and accuracy over large areas that satisfy the requirements of broad range of applications. Such maps are expected to provide highly significant improvements compared to existing products, which tend to be coarser in resolution, and often fail to capture fragmented small-holder farms especially in regions with high dynamic change within and across years. To overcome these limitations, in this research we present an approach for cropland extent mapping at high spatial resolution (30-m or better) using the 10-day, 10 to 20-m, Sentinel-2 data in combination with 16-day, 30-m, Landsat-8 data on Google Earth Engine (GEE). First, nominal 30-m resolution satellite imagery composites were created from 36,924 scenes of Sentinel-2 and Landsat-8 images for the entire African continent in 2015-2016. These composites were generated using a median-mosaic of five bands (blue, green, red, near-infrared, NDVI) during each of the two periods (period 1: January-June 2016 and period 2: July-December 2015) plus a 30-m slope layer derived from the Shuttle Radar Topographic Mission (SRTM) elevation dataset. Second, we selected Cropland/Non-cropland training samples (sample size 9791) from various sources in GEE to create pixel-based classifications. As supervised classification algorithm, Random Forest (RF) was used as the primary classifier because of its efficiency, and when over-fitting issues of RF happened due to the noise of input training data, Support Vector Machine (SVM) was applied to compensate for such defects in specific areas. Third, the Recursive Hierarchical Segmentation (RHSeg) algorithm was employed to generate an object-oriented segmentation layer based on spectral and spatial properties from the same input data. This layer was merged with the pixel-based classification to improve segmentation accuracy. Accuracies of the merged 30-m crop extent product were computed using an error matrix approach in which 1754 independent validation samples were used. In addition, a comparison was performed with other available cropland maps as well as with LULC maps to show spatial similarity. Finally, the cropland area results derived from the map were compared with UN FAO statistics. The independent accuracy assessment showed a weighted overall accuracy of 94, with a producers accuracy of 85.9 (or omission error of 14.1), and users accuracy of 68.5 (commission error of 31.5) for the cropland class. The total net cropland area (TNCA) of Africa was estimated as 313 Mha for the nominal year 2015.

Cropland mapping; cropland areas; 30-m; Landsat-8;↗