Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “manual curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Data Albums: An Event Driven Search, Aggregation and Curation Tool for Earth Science

One of the largest continuing challenges in any Earth science investigation is the discovery and access of useful science content from the increasingly large volumes of Earth science data and related information available. Approaches used in Earth science research such as case study analysis and climatology studies involve gathering discovering and gathering diverse data sets and information to support the research goals. Research based on case studies involves a detailed description of specific weather events using data from different sources, to characterize physical processes in play for a specific event. Climatology-based research tends to focus on the representativeness of a given event, by studying the characteristics and distribution of a large number of events. This allows researchers to generalize characteristics such as spatio-temporal distribution, intensity, annual cycle, duration, etc. To gather relevant data and information for case studies and climatology analysis is both tedious and time consuming. Current Earth science data systems are designed with the assumption that researchers access data primarily by instrument or geophysical parameter. Those who know exactly the datasets of interest can obtain the specific files they need using these systems. However, in cases where researchers are interested in studying a significant event, they have to manually assemble a variety of datasets relevant to it by searching the different distributed data systems. In these cases, a search process needs to be organized around the event rather than observing instruments. In addition, the existing data systems assume users have sufficient knowledge regarding the domain vocabulary to be able to effectively utilize their catalogs. These systems do not support new or interdisciplinary researchers who may be unfamiliar with the domain terminology. This paper presents a specialized search, aggregation and curation tool for Earth science to address these existing challenges. The search tool automatically creates curated "Data Albums", aggregated collections of information related to a specific science topic or event, containing links to relevant data files (granules) from different instruments; tools and services for visualization and analysis; and information about the event contained in news reports, images or videos to supplement research analysis. Curation in the tool is driven via an ontology based relevancy ranking algorithm to filter out non-relevant information and data.

Ramachandran, Rahul↗

Computer Vision Pipeline for Image Analysis for Freeze‐Fracture Electron Microscopy: Rosette Cellulose Synthase Complexes Case

In materials science, plant biology, agriculture, and environmental research, the automated analysis of high-magnification, complex microscopy images, such as those generated by freeze-fracture electron microscopy (FF-TEM), remains a critical challenge that limits the scalability of data interpretation. We present a deep learning computer vision pipeline for high-throughput detection and morphological characterization analysis of cellulose synthase complexes (CSCs, or rosettes) in FF-TEM images. The pipeline integrates preprocessing, detection, human-in-the-loop verification, and semantic segmentation to quantify features such as rosette diameter and inter-lobe spacing. The approach was trained and tested on a curated dataset of high-resolution FF-TEM micrographs of Physcomitrium patens, expanded via strategic tiling and augmentation to over 650 images. We compare YOLOv8 and YOLOv9 architectures and demonstrate that YOLOv9 achieves superior performance in both localization accuracy (mAP50-95 = 0.854) and inference speed. The resulting distributions revealed biological variability consistent with prior manual studies, validating the approach for high-throughput applications. Our results show that the pipeline achieves human-expert level accuracy while dramatically reducing analysis time, enabling scalable, reproducible structural characterization of intramembrane protein complexes. The pipeline is broadly applicable to other domains requiring precise interpretation of complex microscopy data and establishes a foundation for future artificial intelligence (AI)-assisted workflows in biological imaging.

59 BASIC BIOLOGICAL SCIENCES↗

BETO 2021 Peer Review - Electrocatalytic CO2 Utilization

The goal of the ChemCatBio DataHub project is to accelerate the catalyst and process development cycle by developing transformational tools for prediction and collaboration in catalyst R&D. The project is currently focused on the development of the Catalyst Property Database (CPD), a free and public resource released in September 2020. The CPD was designed to advance the state of the art for application of computational data. When computational data, such as computed reaction energetics, is used in catalyst design, it is almost always generated by the researchers seeking to use it, even if similar data has been published previously. One barrier to data reuse that results in this duplication of effort is the difficult process of finding and applying published data, which can be slow, error-prone, and manual. The CPD seeks to overcome these challenges by creating a centralized, searchable database of quality catalyst property data. At present, the CPD contains computed adsorption energies for intermediates along catalytic pathways. During FY21 and FY22, development of the CPD continues with a focus on external users and meeting their requirements. A batch upload capability, training and curation procedures, user interviews, and a demonstration of the CPD's utility in accelerating catalyst research are planned. Overall, the Data Hub project and CPD aim to reduce the time and cost of catalyst research by harnessing the power of data in catalyst discovery.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

BETO 2021 Peer Review - ChemCatBio Data Hub 2.6.2.500

The goal of the ChemCatBio DataHub project is to accelerate the catalyst and process development cycle by developing transformational tools for prediction and collaboration in catalyst R&D. The project is currently focused on the development of the Catalyst Property Database (CPD), a free and public resource released in September 2020. The CPD was designed in response to the observation that when data, such as computed reaction energetics, is used in catalyst design, it is almost always generated by the researchers seeking to use it, even if similar data has been published previously. One barrier to data reuse that results in this duplication of effort is the difficult process of finding and applying published data, which can be slow, error-prone, and manual. The CPD seeks to overcome these challenges by creating a centralized, searchable database of quality catalyst property data. At present, the CPD contains computed adsorption energies for intermediates along catalytic pathways. During FY21 and FY22, development of the CPD continues with a focus on external users and meeting their requirements. A batch upload capability, training and curation procedures, user interviews, and a demonstration of the CPD's utility in accelerating catalyst research are planned. Overall, the Data Hub project and CPD aim to reduce the time and cost of catalyst research by harnessing the power of data in catalyst discovery.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)↗

Leveraging unlabeled SEM datasets with self-supervised learning for enhanced particle segmentation

Scanning Electron Microscopes (SEMs) are widely used in experimental science laboratories, often requiring cumbersome and repetitive user analysis. Automating SEM image analysis processes is highly desirable to address this challenge. In particle sample analysis, Machine Learning (ML) has emerged as the most effective approach for particle segmentation. However, the time-intensive process of manually annotating thousands of SEM images limits the applicability of supervised learning approaches. Self-Supervised Learning (SSL) offers a promising alternative by enabling knowledge extraction from raw, unlabeled data. This study presents a framework for evaluating SSL techniques in SEM image analysis, focusing on novel methods leveraging the ConvNeXtV2 architecture for particle detection. A dataset comprising 25,000 SEM images is curated to benchmark these proposed SSL methods. The results demonstrate that ConvNeXtV2 models, with varying parameter counts, consistently outperform other techniques in particle detection across different length scales, achieving up to a 34% reduction in relative error compared to established SSL methods. Furthermore, an ablation study explores the relationship between dataset size and SSL performance, providing actionable insights for practitioners regarding model selection and resource efficiency. This research advances the integration of SSL into autonomous analysis pipelines and supports its application in accelerating materials science discovery.

Rettenberger, Luca↗

Applicability of Micro X-Ray Fluorescence Spectroscopy to Astromaterials Curation and Research

Introduction: The Astromaterials Acquisition and Curation Office at NASA’s Johnson Space Center (JSC) curates NASA’s astromaterial sample collections which includes: Apollo samples, Luna samples, Ant-arctic meteorites, cosmic dust particles, microparticle impacts into space-flown materials, Genesis solar wind atoms, Stardust comet Wild-2 particles, Stardust inter-stellar particles, Hayabusa asteroid Itokawa particles, Hayabusa 2 asteroid Ryugu particles, and future OSIRIS-Rex asteroid Bennu particles (landing in Sep-tember, 2023) [1–3]. To enhance JSC’s advanced cu-ration capabilities, we have recently installed a high-performance micro-X-ray fluorescence (µXRF) spec-trometer to assist in sample characterization through rapid, non-destructive, in-situ elemental analyses that do not require the sample preparation protocols (i.e., polishing and carbon-coating) commonly needed for electron beam analyses. With this new instrument, we are capable of detecting all elements down to carbon in a variable-pressure or He-purged chamber for anal-ysis of a wide range of sample types. Here we describe the instrumental set-up, capabilities, and applicability of µXRF analysis to astromaterials curation and re-search. Instrumentation and Methodology: The X-ray fluorescence and computed tomography lab (X-FaCT) lab at JSC is now equipped with a Bruker M4 Tornado Plus µXRF (Fig. 1). This system is an energy-dispersive x-ray spectrometer equipped with two 60 mm2 silicon drift detectors (SDD) that are able to be used simulta-neously for output count rates ~500,000 cps. New light element windows allow detecting and analyzing the entire elemental range from carbon to americium. Two x-ray tubes (micro-focus Rh with polycapillary lenses and W with collimators of 0.5, 1.0, 2.0, and 4.5 mm) with max excitation parameters of 50 kV, 30 W and 50 kV, 40 W, respectively, allow for more flexibility of the analysis of high energy lines. The motorized X-Y-Z stage has a mapping range of 190 x 160 mm and can support samples up to 7 kg (~15.5 lbs) and a height of 120 mm [4]. Analytical modes include elemental analysis (down to ~20 µm spot size) via point, line, or area of bulk materials (rock surfaces, thin sections, thick sections, etc.) as well as coating analysis (determination of thickness and composition) of samples. This system has a variable vacuum chamber (1 mbar to 1 atm) that is also equipped with a He-purge system which accommodates vacuum sensitive samples while still allowing detection of light elements at atmospheric pressure. Utility and Applicability of µXRF in Astro-materials research and exploration science (ARES): Elemental analysis using µXRF is commonly em-ployed for both terrestrial and planetary geological science disciplines [5]. It is especially useful for analy-sis of astromaterials given the limited sample prepara-tion required, which is not feasible for certain materi-als. Here we show select applications of µXRF anal-yses of astromaterials that can, have, and will be done at JSC’s X-FaCT lab. Point analysis: In-situ spot analyses (~20 µm spot size) on a cut slab of Martian meteorite NWA 10922 allowed for the discovery, qualitative elemental analy-sis and determination of different feldspar minerals [6]. These point analyses served as an effective prelim-inary step for subsequent quantitative analyses. Ana-lytical standards can be employed for more accurate quantification of µXRF spot analyses. Area analysis: This analytical mode measures all detectable elements (from C to Am) at each pixel (>5 µm pixel size) in a user-defined area. The results are shown as elemental maps which can be extracted as 16-bit TIFF’s for further data processing. In Fig. 2. we show elemental distribution maps of the high-Ti basalt 73001,531 that have been processed using ImageJ software. From these maps you can accurately and quickly (this map took ~50 mins.) identify mineral components, such as pyroxene, plagioclase, oxides, and phosphates, compositional zoning, and mineral textures. Detection of high-Z phases: µXRF techniques are es-pecially effective at analyzing trace minerals with high-atomic-number (high-Z) elements because the high-energy characteristic X-rays used (relative to SEM EDS) allow for mapping of K lines in elements up to La (typically SEM maps use L X-ray lines for elements >Zn, and these can often have interferences). Thus, µXRF is especially suited for identifying minerals like zircon, baddeleyite, REE-rich phosphates, Fe-rich met-als, oxides, sulfides, and phosphides [4]. In Fig. 3 we show elemental distribution maps for 73001,530 where we are able to correlate the original video image with, Zr, Si, Y and Hf elemental maps together identi-fying the location of a zircon. In this location you would expect lower Si compared to surrounding mate-rial, as well as higher Zr, Y, and Hf content compared to surrounding material, all of which is confirmed by our XRF ele-mental distribution maps (Figure 2.) Conclusions: The new M4 Tornado Plus µXRF within the Astromaterials Acquisition and Curation office at NASA JSC allows for rapid and non-destructive elemental analysis of astromaterials with limited or no sample preparation. µXRF analyses pro-vide crucial compositional knowledge for the prelimi-nary examination and curation of astromaterials. This instrument enhances the advanced curation capabili-ties in the X-FaCT laboratory at JSC by allowing pro-ductive, cohesive, and non-destructive multi-modal x-ray analyses on astromaterial samples, which is neces-sary for the comprehensive curation and study of our current and future astromaterial collections. Addition-ally, µXRF can provide complimentary information to researchers for studies on astromaterials. References: [1] Allen, C. et al., (2011). Chemie De Erde Geochemistry, 71, 1-20. [2] McCubbin, F. M. et al., (2016) 47th LPSC, abstract #2668 [3] Zeigler, R. A. et al., (2017) 48th LPSC, abstract #2772 [4] Bruker User Manual [5] Young et al., (2016) Appl. Geochemistry, 72, 77-87 [6] Mor-ris, R. V. et al., (2023) 54th LPSC.

E W O'Neal↗

DOE COVID-19 Data Curation Effort: Overview of Initial Data Collection Coverage (March - June 2020)

During the COVID-19 pandemic of 2020, major case reporting outlets quickly coalesced around two or three primary vendors. Johns Hopkins University and The New York Times were among the more prominent, and all were of great value to the nation, particularly during the uncertain early stages of the pandemic. They primarily focused on three major attributes: number of new cases, deaths, and recovery, but only at the state level. Recognizing that many states were reporting very detailed data sets (e.g., hospital beds) at a count level or finer, the ORNL Pandemic Modeling team embarked on a major data curation effort from March to June 2020 for the purpose of capturing this wealth of detailed data. The challenge of curating this data was daunting. The number of attributes reported by the states grew on almost on a weekly basis. States were routinely shifting their web tool strategies away from easily parsable HTML-based formatting to new Tableau and ArcGIS content. This growth in the sheer number of attributes combined with the unpredictable shifts in data format meant an aggressive and agile combination of automated scripting and manual scraping was required to capture new daily streams. To keep up, the team had to scale up staff and widen its approach for capture and storage. The DOE COVID-19 data collection effort resulted in over 11 million data points being collected, covering over 13,000 unique geographies and over 2,000 unique attributes that spanned predominantly from early March through the end of June 2020.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)↗

Scientific Core Library Stack (SCLS) v2026

SCLS (Scientific Core Library Stack) is an opinionated build and packaging system for scientific computing libraries developed at Lawrence Berkeley National Laboratory. It produces a coherent, reproducible stack of numerical libraries — including BLAS/LAPACK, MPI, sparse direct and iterative solvers, graph partitioners, and parallel I/O libraries (e.g., PETSc, SLEPc, HDF5, NetCDF, MUMPS, OpenBLAS) — that work together without manual repair by downstream scientific software. From a single recipe-and-flavor model, SCLS produces native RPM packages for RHEL-family Linux, DEB packages for Debian/Ubuntu, direct Unix-style prefix installs for HPC and locked-down environments, and native macOS builds. Multiple build "flavors" (e.g., GCC+OpenBLAS, GCC+MKL, Intel+MKL, debug) coexist in distinct prefixes on the same host. Compared to general-purpose meta-build frameworks, SCLS is deliberately curated rather than infinitely configurable. It enforces deterministic, audit-friendly behavior: explicit build dependencies, no silent feature autodetection, a clear open-source license policy, and rpath-based runtime linkage so installs integrate cleanly with standard package-manager workflows.

Messe, Christian [Lawrence Berkeley National Labor↗

In-Lab Rapid Analytical Detection of Lunar Volatiles By Universal Gas Analyzer With Comparison to GC-MS System

Introduction: The curation of permanently shadowed regions (PSRs) [1] on the lunar surface is centered around studies based upon the observed volatiles from the LCROSS mission [2]. The rapid detection of important volatile gases and vapors present in planetary bodies and Astromaterials by a standalone analytical device is an area of intense research interest in our group and Planetary Exploration & Astromaterials Research Laboratory (PEARL) facility and this work is relevant to the future preparation of viable lunar simulants for testing curation efforts down the road. The groundbreaking results obtained from the LCROSS Mission [2] open the requirements for the direct detection of volatiles present in regolith materials collected from the lunar surface. The mass spectrometry of volatile chemicals is a general technique that utilizes a set of instruments that creates charged ions from a gaseous chemical species and measures the intensities vs. mass-to-charge ratio (m/z) [3]. In this context, we discuss in-lab experimental results and procedures for rapid qualitative analysis of main LCROSS volatiles (water, H2S, NH3, CO2, and CH3OH) by a Universal Gas Analyzer (UGA) instrument. Additionally, the instrument performance was evaluated by measuring the isotopic abundance ratio of atmospheric Ar-40 to Ar-36 present in room air since, argon is a relevant gas in planetary studies as it can provide an insight and reference point to isotope studies [4]. Additional, cross comparisons were attempted and made between the two instruments to develop a robust analytical technique by comparing mass spectral data for H2S headspace samples with a Trace-1310/ ISQ 7000 (ThermoFisher Scientific.) GC-MS system. Background: The benchtop UGA System is equipped with an SRS UGA 300 quadrupole mass spectrometer designed and built by Stanford Research Systems [5]. This system can be configured for several types of gaseous chemical analysis. The inlet line continuously samples gases at low flow rates (several milliliters per minute) through a capillary limiting the intake pressure making the instrument ideal for online analysis of select gases and/or room atmosphere. Moreover, in our current UGA system, a change in composition at the inlet can be detected in about 200 milliseconds and a complete spectrum is acquired (for a range of 1-100 amu) in under 45 sec with masses measured at rates up to 25 msec per point [5]. This system provides a quick upstream analytical data that we can then compare to results obtained by our GC-MS system. Sample Preparation: Small volume (2-4 mL) of analyte sample was taken in a 10 mL glass vial and sealed with a crimped cap and purged with pure Ar or N2 gas to displace air from the top. The headspace sample was scanned by the UGA instrument at analog, histogram, and pressure vs. time modes. The isotopic abundance ratio for 40Ar-to-36Ar was estimated by measuring partial pressure vs. time scans and setting the mass at 40 and 36 respectively. Results and Discussions: In this work, we have investigated the applicability of the UGA system by qualitative analysis of a series of LCROSS volatiles measured individually. Fig. 1 demonstrates a set of vertically offset spectra for the partial pressures measured as a function of mass-to-charge (m/z) ratios. The average acquisition time for each spectrum was less than a minute suggesting that the UGA system is ideal for quick analysis of geochemical volatiles. For the cross-comparison, we analyzed an H2S headspace sample by a Trace-1310/ISQ-7000 system and compared mass spectral data with previously measured UGA histogram scan data (Fig. 2). In both cases, major peak positions are the same, however, the intensities of fragment ions ([1H132S]+ and [32S]+) are higher for UGA suggesting that the fragment ionization process is stronger in UGA compared to that of GC-MS. To investigate how the integrated area under each chromatogram varies with the headspace sample volume, a set of five H2S headspace samples with increasing volumes was analyzed by the GC-MS system (Fig. 3, inset). A small volume (e.g., 200 to 1000 µL) of H2S/H2O vapor was withdrawn from a 20 mL stock sample vial containing ~5 mL of 0.4% H2S in water by a gas-tight syringe and added to another 20 mL vial filled with argon and analyzed by the GC-MS system. Finally, the UGA detector sensitivity was evaluated by calculating the atmospheric 40Ar-to-36Ar isotopic abundance ratio in room air by running a partial pressure vs. time scan with setting the atomic mass at 40, and 36. Fig. 4(a) shows a ~25 min duration “P vs. time” scan for 40Ar (plot for 36Ar is not shown). The partial pressure values (~100 points) were corrected by subtracting the corresponding background pressure value for 37Ar and utilized to calculate 40Ar-to-36Ar isotopic abundance ratios as shown by Fig. 4b. The average isotopic abundance ratio is ~306 with a 2*STDEV ~13. This abundance ratio is significantly close to the previously reported value of 298.56 [6] and the ratio obtained by our GC-MS system (303 for a UHP grade Ar sample). Conclusions: Our study strongly evidenced that the benchtop UGA system is a valuable analytical tool for the detection of major LCROSS volatiles. The rapid scanning capability, the inexpensiveness of the whole system, and impressive detection sensitivity prove its worthiness as an essential device for advanced geochemical applications. Moreover, cross comparisons with the GC-MS provide important bridges into advanced curatorial efforts into the future. References: [1] Bickel, V.T., et al. (2021) Nat Commun 12, 5607. [2] Colaprete, A., et al. (2010) Science, 330, 463-468. [3] Glavin, D. P. et al. (2012) 2012 IEEE Aerospace Conference, 1-11. [4] Willett, C. D., et al. (2022) Geochimica et Cosmochimica Acta 329, 119-134. [5] Operation Manual and Programming Reference. (2018) Universal gas Analyzers, Stanford Research Systems. [6] Lee, J. Y., et al. (2006) Geochimica et Cosmochimica Acta 70, 4507–4512. Notes: (4 figures are attached with text as shown by the attached file)

Curation↗

An Open Combinatorial Diffraction Dataset Including Consensus Human and Machine Learning Labels with Quantified Uncertainty for Training New Machine Learning Models

Modern machine learning and autonomous experimentation schemes in materials science rely on accurate analysis of the data ingested by these models. Unfortunately, accurate analysis of the underlying data can be difficult, even for domain experts, complicating the training of the models intended to drive experiments. This is especially true when the goal is to identify the presence of weak signatures in diffraction or spectroscopic datasets. In this work, we examine a set of as-obtained diffraction data that track the phase transition from monoclinic to tetragonal in a Nb-doped VO2 film as a function of temperature and dopant concentration. We then task a set of domain experts and a set of machine learning experts with identifying which phase is present in each diffraction pattern manually and algorithmically, respectively; in both cases, the labels can vary dramatically, especially at the phase boundaries. We use the mode of the labels and the Shannon entropy as a method to capture, preserve and propagate consensus labels and their variance. Further we use the expert labels as a benchmark and demonstrate the use of Shannon entropy weighted scoring to test the performance of machine learning generated labels. Finally, we propose a material data challenge centered around generating improved labeling algorithms. This real-world dataset curated with expert labels can act as test bed for new algorithms. The raw data, annotations and code used in this study are all available online at data.gov and the interested reader is encouraged to replicate and improve the existing models

97 MATHEMATICS AND COMPUTING↗

Identifying Disinformation Using Rhetorical Devices in Natural Language Models

Foreign disinformation campaigns are strategically organized, extended efforts using disinformation – false or misleading information deliberately placed by an adversary – to achieve some goal. Disinformation campaigns pose severe threats to our nation’s security by misinforming decision makers and negatively influencing their actions when they are operating on limited amounts of evidence. Current efforts rely on subject matter experts to manually identify disinformation, or on computers and traditional natural language processing algorithms to identify patterns in data to calculate the probability that something is disinformation or not. While both have their merits and successes, subject matter experts are unable to keep up with the high volumes of global information and traditional natural language algorithms do not do well in identifying “why” something is disinformation or not. Our hypothesis is that we can identify disinformation by looking at the way someone speaks, in the rhetorical devices they use. We have curated and annotated a dataset designed for multiple natural language processing tasks, but specifically useful for disinformation detection algorithms.

97 MATHEMATICS AND COMPUTING↗

A Multimodal Event Catalog and Waveform Data Set That Supports Explosion Monitoring from Nevada, U.S.A.

Multimodal, curated data sets and nuisance event catalogs remain rare in the explosion monitoring community relative to curated seismic data sets. The source of this relative absence is the difficultly in deploying multimodal receivers that sense the seismic, acoustic, and other modalities from multiphysics sources. We provide such a data set in this study that delivers seismic, infrasound, and electromagnetic (magnetometer) sensor records collected over a two–week period, within 255 km of a 10 ton buried chemical explosion called DAG–4 that was located at 37.1146°, –116.0693° on 22 June 2019 21:06:19.88 UTC. This catalog includes 485 seismic, seismoacoustic, and infrasound–only events that an expert analyst manually built by reviewing waveforms from 29 seismic and infrasound sensors. Our data release includes waveforms from these 29 seismic, infrasound, and seismoacoustic stations and two magnetometer stations and their station metadata. We deliver these waveforms in NNSA KB Core CSS.w format (i4) with a corresponding wfdisc table that provides the header information. Here, we expect that this data set will provide a valuable, benchmark resource to develop signal processing algorithms and explosion monitoring methods against manual, human observations.

58 GEOSCIENCES↗

Natural Language Processing for Text Based Event Extraction: Identifying Events of Interest Related to Worldwide State-Sponsored Civil Nuclear Power

Beginning in FY20, SRNL was funded by the National Nuclear Security Administration’s Office of Defense Nuclear Non-Proliferation Research and Development to develop a prototype natural language processing/natural language understating machine learning-based modeling and analysis pipeline to extract and forecast events of interest from massive open data sources. The working hypothesis within the approach is that contextual shifts in key words and phrases act as indicators of events of interest over time. Therefore, by identifying points in time where contextual shifts occur, events of interest can be extracted along with explicit and implicit connections of entities and activities. The development of the preliminary prototype pipeline proved successful, meriting further testing of the pipeline on more broad topical domains and in a worldwide data environment. Therefore, SRNL, in collaboration with the Sanghani Center for Artificial Intelligence and Data Analytics at Virginia Tech, have continued development with a test case of identifying events of interest related to worldwide state-sponsored civil nuclear power in open data sources. In the first year of this follow-on effort, the team has curated domain-specific data corpuses using an automated scheme and applied the modeling and analysis pipeline. This robust, focused, and efficient approach consists of an ensemble of analyses applied to time dependent word embedding models that are trained on the data corpuses. In this report, the team has demonstrated the capability of the existing pipeline (as development has continued in parallel) by exploring several specific case-studies centered around Rosatom’s international activities regarding the planning, construction, operation, and/or shutdown of nuclear reactors. A basic timeline events has been generated by manually cataloging known “milestone” events that have occurred at reactors in Turkey, Finland, Hungary, and Egypt and compared with the output of the modeling pipeline. In this approach, the team has characterized the lead time using the prototype pipeline, as well as the ability to capture relevant information, which proved 100% successful. A deep dive example of the Akkuyu reactor (Turkey) is presented that shows the breadth of information that can be captured using the approach. In this case study, events were extracted pertaining to the planning/construction of Akkuyu including protests from the population, information campaigns in response to the protests, forged regulatory documents and lawsuits, budgetary/shareholder information, geopolitical tensions, and the various construction milestones. This has demonstrated the pipeline’s utility as a research aid or real-time event extraction tool, where summary-level information and detailed text extractions from millions of articles or Tweets across long time periods can be generated with significantly less effort than current techniques.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Distribution Substation Planning Toolkit (dsp-toolkit) v1.0

The Distribution Substation Planning Toolkit (DSP Toolkit) is a software suite designed to streamline the planning and optimization of distribution substations. This toolkit offers a comprehensive set of tools and APIs for data curation, short-term electric load forecasting, and weather-sensitive load adjustment, making it an essential resource for utility companies, engineers, and researchers. Features • Data Preprocessing and Curation: Efficiently manage and preprocess large datasets to ensure high-quality input for analysis. • Short-Term Load Forecasting: Utilize data-driven models to predict short-term electric loads accurately. • Weather-Sensitive Modeling: Automatically adjust load forecasts based on weather data to predict future peak demands more precisely. Uses The DSP Toolkit is ideal for planning and optimizing distribution substations, providing a user-friendly interface and comprehensive documentation. It is suitable for both novice and experienced users, facilitating efficient and accurate planning processes. Advantages • Efficiency: Automates complex planning tasks, reducing manual effort and minimizing errors. • Scalability: Handles large datasets and complex models, making it suitable for large-scale projects. • Community and Support: Open-source with active community contributions, ensuring continuous improvement and support. • Extensibility: Easily extendable with custom modules and plugins, allowing users to tailor the toolkit to their specific needs. The DSP Toolkit stands out by offering a robust, flexible, and user-friendly solution for distribution substation planning. Public Abstract

Li, Han [Lawrence Berkeley National Laboratory (LB↗

Deep Point Cloud Building Envelope Segmentation (DeeP-CuBES) using Deep Learning

Building Information Modeling (BIM) plays an important role in building design and construction, particularly for achieving energy-efficient retrofits. Building envelope retrofits using panelized prefabricated system, such as those popularized by the Energiesprong program, need accurate as-built dimensions of facade features (windows, doors, etc.) to achieve the desired thermal and air tightness. Traditionally, building surveying is done manually, resulting in a time-consuming and labor-intensive process. Recently, 3D point clouds from terrestrial LiDAR have been used to automate the generation of as-built dimensions of existing buildings. However, automated BIM using LiDAR relies on solving the point cloud semantic segmentation (PCSS) problem. In this work, we propose a robust pipeline for solving the PCSS problem using deep neural networks, focusing on overcoming challenges posed by imbalanced datasets and complex architectural features. We introduce the first high-density, labeled, and validated building envelope point cloud dataset derived from multiple building scans, specifically curated to tackle challenges in facade-level segmentation. Results from the trained neural networks show that advanced attention-based architectures and incorporating radiometry (light intensity and RGB) features significantly boost segmentation accuracy for windows and doors.

Selvakumar, Balaji [ORNL]↗

Image analysis used to count and measure etched tracks from ionizing radiation

We have developed techniques to use digitized scanning electron micrographs and computer image analysis programs to measure track densities in lunar soil grains and plastic dosimeters. Tracks in lunar samples are formed by highly ionizing solar energetic particles and cosmic rays during near surface exposure on the Moon. The track densities are related to the exposure conditions (depth and time). Distributions of the number of grains as a function of their track densities can reveal the modality of soil maturation. We worked on two samples identified for a consortium study of lunar weathering effects, 61221 and 67701. They were prepared by the lunar curator's staff as polished grain mounts that were etched in boiling 1 N NaOH for 6 h to reveal tracks. We determined that backscattered electron images taken at 10 percent contrast and approximately 50 percent brightness produced suitable high contrast images for analysis. We used the NIH Image program to cut out areas that were unsuitable for measurement such as edges, cracks, etc. We ascertained a gray-scale threshold of 25 to separate tracks from background. We used the computer to count everything that was two pixels or greater in size and to measure the area to obtain track densities. We found an excellent correlation with manual measurements for track densities below 1 x 10(exp 8) cm(exp -2). For track densities between 1 x 10(exp 8) cm(exp -2) to 1 x 10(exp 9) cm(exp -2) we found that a regression formula using the percentage area covered by tracks gave good agreement with manual measurements. We determined the track density distributions for 61221 and 67701. Sample 61221 is an immature sample, but not pristine. Sample 67701 is a submature sample that is very close to being fully mature. Because only 10 percent of the grains have track densities less than 10(exp 9) cm(exp -2), it is difficulty to determine whether the sample matured in situ or is a mixture of a mature and a submature soil. Although our analysis of plastic dosimeters is at an early stage of development, results are encouraging. The dosimeter was etched in 6.25 N NaOH at 70 deg C for 16 h. We took 200x secondary electron images of the sample and used the NIH Image software to count and measure major and minor diameters of the etched tracks. We calculated the relative track etch rate from a formula that relates it to the major and minor diameters. We made a histogram of the number of tracks versus their relative etch rate. The relative track etching rate is proportional to the linear energy transfer of the particle. With appropriate calibration experiments, the histogram could be used to calculate the radiation dose.

Blanford, George E.↗