Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Subject-specific modeling framework for particle deposition using computational fluid dynamics

Quantifying particle deposition and dose in the respiratory tract requires a physiologically realistic representation and reproducible computational workflows. However, existing modeling frameworks, such as the International Commission on Radiological Protection (ICRP) compartmental models and the Multiple Path Particle Dosimetry (MPPD) tool, lack detailed deposition profiles and subject-specific capabilities. The combination of advances in computer vision algorithms applied to the respiratory tract and Computational Fluid and Particle Dynamics (CFPD) allows high-fidelity simulations of particle behavior in anatomically accurate geometries derived from individual CT scans. The segmentation, preprocessing, and file preparation task for a CFPD simulation was often time-consuming, and no prior studies to-date have yet presented a fully automated framework. This work presents a fully automated workflow to obtain individualized particle deposition profiles in the human respiratory tract. The pipeline starts with segmenting upper and lower airway geometries using morphological and deep learning-based methods, generating three-dimensional (3D) models from CT imaging data. Next, a series of algorithms are presented to quality check and prepare the 3D geometry for a CFD or CFPD simulation. The preprocessing step includes correcting geometric artifacts, enforcing a physically consistent mesh, and automatically identifying and capping multiple outlets, which is required for CFD/CFPD simulations. These processed models are then input into open-source (OpenFOAM) or commercial (StarCCM+) CFD solvers, where flow and transient particle transport equations — including turbulence and particle–wall interactions are solved under realistic breathing conditions. Finally, the resulting particle deposition profiles can be integrated with Monte Carlo radiation transport codes and state-of-the-art computational phantoms to assess organ-specific absorbed doses in scenarios of radioactive aerosol inhalation. The presented work streamlines respiratory tract segmentation, preprocessing for CFD/CFPD simulations, and integration with dose assessment workflows, reducing manual intervention and improving access to high-fidelity, subject-specific modeling. The high precision in predicted particle deposition and dose distributions can improve personalized treatment strategies in respiratory medicine and refine dose estimates for radiation protection.

AI↗

Data Science Techniques, Assumptions, and Challenges in Alloy Clustering and Property Prediction

Data analytics methods have been increasingly applied to understanding materials chemistry, processing due to the manufacturing approach, and uni-axial and cyclic property relationships in the highly complex space of alloy design. There are several benefits to applying data analytics to this space, including the ability to manage non-linearities in the responses of the alloy attributes and the resulting mechanical properties. However, key difficulties in applying and understanding the results of data analytics include the often lack of reported assumptions and data processing steps necessary to improve interpretation and reproducibility in derived results. In this work, the methods used to generate clustering and correlation analyses for experimental 9% Cr ferritic-martensitic steel data were investigated and the resulting implications for mechanical property predictions were assessed. This work uses principal component analysis, partitioning around medoids, t-SNE, and k-means clustering to investigate trends in composition, processing and microstructure information with creep and tensile properties, building on work done previously using a smaller version of the same dataset. The initial assumptions, preprocessing steps and methods are investigated and outlined in order to depict the fine level of detail required to convey the steps taken to process data and produce analytical results. Here, the variations in the resulting analyses are explored due to the influence of new and more varied data.

36 MATERIALS SCIENCE↗

Diffractive optical computing in free space

Abstract Structured optical materials create new computing paradigms using photons, with transformative impact on various fields, including machine learning, computer vision, imaging, telecommunications, and sensing. This Perspective sheds light on the potential of free-space optical systems based on engineered surfaces for advancing optical computing. Manipulating light in unprecedented ways, emerging structured surfaces enable all-optical implementation of various mathematical functions and machine learning tasks. Diffractive networks, in particular, bring deep-learning principles into the design and operation of free-space optical systems to create new functionalities. Metasurfaces consisting of deeply subwavelength units are achieving exotic optical responses that provide independent control over different properties of light and can bring major advances in computational throughput and data-transfer bandwidth of free-space optical processors. Unlike integrated photonics-based optoelectronic systems that demand preprocessed inputs, free-space optical processors have direct access to all the optical degrees of freedom that carry information about an input scene/object without needing digital recovery or preprocessing of information. To realize the full potential of free-space optical computing architectures, diffractive surfaces and metasurfaces need to advance symbiotically and co-evolve in their designs, 3D fabrication/integration, cascadability, and computing accuracy to serve the needs of next-generation machine vision, computational imaging, mathematical computing, and telecommunication technologies.

36 MATERIALS SCIENCE↗

Statistical characterization of experimental magnetized liner inertial fusion stagnation images using deep-learning-based fuel–background segmentation

Significant variety is observed in spherical crystal x-ray imager (SCXI) data for the stagnated fuel–liner system created in Magnetized Liner Inertial Fusion (MagLIF) experiments conducted at the Sandia National Laboratories Z-facility. As a result, image analysis tasks involving, e.g., region-of-interest selection (i.e. segmentation), background subtraction and image registration have generally required tedious manual treatment leading to increased risk of irreproducibility, lack of uncertainty quantification and smaller-scale studies using only a fraction of available data. We present a convolutional neural network (CNN)-based pipeline to automate much of the image processing workflow. This tool enabled batch preprocessing of an ensemble of N scans = 139 SCXI images across N exp = 67 different experiments for subsequent study. The pipeline begins by segmenting images into the stagnated fuel and background using a CNN trained on synthetic images generated from a geometric model of a physical three-dimensional plasma. The resulting segmentation allows for a rules-based registration. Our approach flexibly handles rarely occurring artifacts through minimal user input and avoids the need for extensive hand labelling and augmentation of our experimental dataset that would be needed to train an end-to-end pipeline. Here we also fit background pixels using low-degree polynomials, and perform a statistical assessment of the background and noise properties over the entire image database. Our results provide a guide for choices made in statistical inference models using stagnation image data and can be applied in the generation of synthetic datasets with realistic choices of noise statistics and background models used for machine learning tasks in MagLIF data analysis. We anticipate that the method may be readily extended to automate other MagLIF stagnation imaging applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Analytical Modeling of Exoplanet Transit Spectroscopy with Dimensional Analysis and Symbolic Regression

Abstract The physical characteristics and atmospheric chemical composition of newly discovered exoplanets are often inferred from their transit spectra, which are obtained from complex numerical models of radiative transfer. Alternatively, simple analytical expressions provide insightful physical intuition into the relevant atmospheric processes. The deep-learning revolution has opened the door for deriving such analytical results directly with a computer algorithm fitting to the data. As a proof of concept, we successfully demonstrate the use of symbolic regression on synthetic data for the transit radii of generic hot-Jupiter exoplanets to derive a corresponding analytical formula. As a preprocessing step, we use dimensional analysis to identify the relevant dimensionless combinations of variables and reduce the number of independent inputs, which improves the performance of the symbolic regression. The dimensional analysis also allowed us to mathematically derive and properly parameterize the most general family of degeneracies among the input atmospheric parameters that affect the characterization of an exoplanet atmosphere through transit spectroscopy.

79 ASTRONOMY AND ASTROPHYSICS↗

DICER: Data Intensive Computing Environment and Runtime for Evaluating Unprecedented Scale of Geospatial-Temporal Human Mobility Data

With the significant increase in sources and volume of human mobility data through commercial data vendors as well as microsimulation of cities, the scale of geospatial-temporal data to analyze and assess for mobility characterization has grown to the level of Big Data. There are mobility related commercial organizations deploying scalable computing, but often the system architecture, workflow, and intermediate processing components are not fully disclosed in relevant scope. Current research literature has a notable lack of studies demonstrating architectures and workflows for human mobility analytics that are implemented on a TeraByte scale of geospatial-temporal data. In this context, this paper presents a hyperscale-level system solution named DICER (Data Intensive Computing Environment and Runtime) for processing and analytics of geospatial-temporal data at big data scale. Although the cluster computing architecture of DICER with Apache Spark job running on Kubernetes cluster is not new, there are innovations in the workflow, hierarchical processing logic, and a wide range of intermediate preprocessing and mobility metrics calculation. We have performed case studies to validate the effectiveness of DICER system solution by performing detailed analytics and assessment of human mobility microsimulation output at three different scopes and scale, including a usecase with 16.97 TeraByte and 259.2 Billion rows of data. In addition, we have presented another case study of utilizing DICER to perform the same mobility processing and comparative analytics on large-scale commercially available geospatial-temporal data. All these case studies validate the efficiency and usefulness of DICER in computing population mobility characteristics from geospatial-temporal trajectory data at an unprecedented scale (not only just data volume, but also combination of: number of user entities, temporal frequency, spatial resolution, data duration).

De, Debraj↗

Leveraging BERT and Network-Based Attention Analysis for Identifying Treatment Milestones in EHRs

This study introduces a sophisticated data-driven framework for analyzing Electronic Health Records (EHRs) using transformer-based models to identify and disentangle overlapping treatment contexts. The framework leverages a preprocessing pipeline that transforms structured procedural codes into semantically enriched descriptive text, enabling the use of attention mechanisms to cluster medical events into treatment milestones—cohesive and distinct components of care processes. The methodology is rigorously validated using synthetic datasets derived from the MIMIC-III database, designed to simulate the heterogeneity and overlapping procedural contexts characteristic of real-world EHR scenarios. Quantitative evaluation highlights the framework’s robustness in disentangling concurrent care pathways, with attention metrics and unsupervised clustering approaches demonstrating the ability to preserve intra-context relationships while distinguishing inter-context dependencies. By addressing challenges inherent in data heterogeneity, this approach provides a foundation for uncovering complex treatment patterns, advancing clinical decision-making, and optimizing resource allocation in diverse healthcare environments.

Kim, Minsu [ORNL] (ORCID:0000000224185535)↗

Expanded analysis of machine learning models for nuclear transient identification using TPOT

Industries around the world are becoming more and more data driven. The nuclear field is no exception with several different applications being proposed. One popular area of research is the use of machine learning in transient detection. This paper seeks to build upon a previous study which made use of the AutoML package TPOT to train traditional machine learning models to classify transient events occurring with a reactor. Synthetic data was once again collected using a GPWR reactor simulator. Data on 12 different events was collected using 15 different initial conditions. Here, a dataset consisting of over 100,000 data points was compiled and used to train 7 different machine learning models using a pre-defined TPOT dictionary with 12 different preprocessing techniques. Three of the trained models were able to produce validation results in the 90s with the expanded dataset. Once the models were trained, it was possible to look into where during the simulation, misclassifications occurred. Using these three models, analysis was done to determine if TPOT could be used to train models that were effective if important features were missing. The results from this were positive with the newly trained models scoring close to the original models. Finally, to conclude this study, the three high performing models were retrained using different random states to see if there was any major variation when different states were used.

42 ENGINEERING↗

The good, the bad, and the ugly: Data-driven load profile discord identification in a large building portfolio

Reducing the overall energy consumption and associated greenhouse gas emissions in the building sector is essential for meeting our future sustainability goals. Recently, smart energy metering facilities have been deployed to enable monitoring of energy consumption data with hourly or subhourly temporal resolution. This unprecedented data collection has created various opportunities for advanced data analytics involving load profiles (e.g., building energy benchmarking programs, building-to-grid integration, and calibration of urban-scale energy models). These applications often need preprocessing steps to detect daily load profile discords, such as: 1) outliers due to system malfunctions (the bad) and 2) irregular energy consumption patterns, such as those resulting from holidays (the ugly) compared to normal consumption patterns (the good). However, current preprocessing methods predominantly focus on filtering using statistical threshold values, which fail to capture the contextual discords of daily profiles. In addition, discord detection algorithms in building research are often aimed at finding individual building-level discords, which are not suitable at a large scale. Thus, here, we develop a method for automated load profile discord identification (ALDI) in a large portfolio of buildings (more than 100 buildings). Specifically, ALDI 1) uses the matrix profile (MP) method to quantify the similarities of daily subsequences in time series meter data, 2) compares daily MP values with typical-day MP distributions using the Kolmogorov-Smirnov test, and 3) identifies daily load profile discords in a large building portfolio. We evaluate ALDI using the metering data of both an academic campus and a residential neighborhood. Our results demonstrate that ALDI efficiently discovers measurement errors by system malfunctions and low energy consumption days in the academic campus portfolio, and it detects unique load shape patterns likely driven by occupant behavior and extreme weather conditions in the residential neighborhood.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Herbaceous Feedstock 2019 State of Technology Report

The U.S. Department of Energy (DOE) promotes the production of advanced liquid transportation fuels from lignocellulosic biomass by funding fundamental and applied research that advances the State of Technology (SOT). As part of its involvement with this mission, Idaho National Laboratory (INL) completes an annual SOT report for biomass feedstock logistics. This report summarizes supply system impacts of Bioenergy Technologies Office (BETO)-funded research and development efforts at INL and elsewhere (such as the High-Tonnage Feedstock Logistics projects (Webb et al. 2013a, Webb et al. 2013b, Webb et al. 2013c, Webb and Sokhansanj 2014, Sokhansanj et al. 2014)) that lead to improvements in feedstock supply systems. These include improvements to and observed performance of innovative harvest and collection methods, storage technologies, transportation and handling approaches, and advanced preprocessing technologies. Biomass quality and variability, and the interface between feedstock quality and conversion performance are key drivers in addition to delivered feedstock cost. In this report, we estimate the benefits of R&D improvements to individual supply system unit operations and present the status of feedstock logistics technology development for converting biomass into biofuels. These analyses are supported by experimental data where possible and help to align the SOT relative to the cost goals defined in the Multi-Year Program Plan. The 2019 Herbaceous SOT incorporates several technology changes in feedstock preprocessing and introduces opportunities from the integrated landscape management (ILM) strategy and increased grower participation to reduce biomass access costs, while maintaining or improving grower profitability. During FY18 uneven flow from the horizontal bale grinder was identified as a significant issue limiting preprocessing system throughput. Based on FSL-funded research at INL, the 2019 Herbaceous SOT replaces the horizontal bale grinder used in the first stage size reduction with a bale processor. The improved uniformity of biomass flow entering the PDU eliminated slugging flow from the first stage size reduction and improved the throughput of downstream operations. In order to achieve moisture reduction through frictional heating during grinding (which allowed elimination of the costly rotary drum dryer in previous SOTs), the second stage grinder was changed from a rotary shear, which does not remove moisture, back to a hammer mill. Finally, the 2019 Herbaceous SOT introduces modified three-pass and two-pass corn stover supply curves derived from the BT16 resource assessment, based on FY19 modeling results (WBS 4.2.1.20) quantifying economic benefits of ILM in the supply area, together with modeling results (WBS 1.2.1.5) identifying ILM strategies to increase grower participation. The 2019 Herbaceous SOT report documents the current modeled cost of an herbaceous feedstock supply system from harvest to the pretreatment reactor throat for hydrocarbon fuel production via biochemical conversion, based on equipment and processes now available or potentially available in the near term. The modeled cost also considers both the required quality and the availability of the biomass resources. The 2019 Herbaceous SOT predicts a modeled delivered feedstock cost of $81.37 /dry ton (2016$); this is a $2.30/dry ton (2016$) decrease from the 2018 Herbaceous SOT. Technology improvements that contributed to this modeled cost reduction include reduced cost for the new preprocessing design and quantification of the opportunities of the integrated landscape management (ILM) strategy and an increased grower participation rate to reduce the grower payment portion of biomass access costs, while maintaining or improving grower profitability. A greenhouse gas emissions (GHG) assessment was completed by Argonne National Laboratory using the 2019 Greenhouse Gases, Regulated Emissions, and Energy use in Transportation model, estimating an increase of 14.89 kg CO2e/ton from the 2018 SOT (69.27 kg CO2e/ton in 2018 to 84.16 kg CO2e/ton in 2019). The increase of energy consumption during preprocessing along with higher transportation distance to access low cost biomass from further distance contributed to the increase of GHG emissions in the 2019 Herbaceous SOT. The reason for the increased transportation distances was the cost tradeoff of going farther from the biorefinery to access the cheaper ILM-derived counties (the cheaper price outweighed the cost of increased supply radius).

09 BIOMASS FUELS↗

HydraGNN v4.0

The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Data Projection of the High Temperature Electrolysis System in the Dynamic Energy Transport and Integration Laboratory using Dynamic System Scaling

For nuclear power to be flexible in a functioning Integrated Energy System (IES), excess produced heat must be stored or utilized during times of low power demand to ensure a load factor of 1 while load balancing. The Dynamic Energy Transport and Integration Laboratory (DETAIL) is one facility that is under development to emulate IES conditions on the engineering-scale, planned to conduct virtual real time operations with industry-scale facilities, and is currently testing thermal storage and high temperature electrolysis. As part of the study to develop a method to preprocess input signals or postprocess output signals between systems of different scales via Dynamical System Scaling (DSS), the current research is one of the continued efforts branching from the data projection activity conducted for the Thermal Energy Distribution System and currently engages the High Temperature Electrolysis (HTE) System in DETAIL. The HTE SOEC electrical, fluid, and thermal dynamics Figure of Merits (FOM) were identified, governing equations and closure relations were successfully scaled, and relations between FOM scaling ratios were determined. Setting the scaling objectives to reform existing data to project a data set that doubly accelerated the electrolysis process while preserving the produced amount of hydrogen was generated for the full transient. The calculated boundary conditions were inlet temperature, stack current, and inlet steam mass flow rate at 1470 K, 121.1 A, and 1.886 g/s, respectively. The research outcomes demonstrated an output signal postprocessing case accelerating the hydrogen production without changing geometry, number of cells, and partial pressures.

08 HYDROGEN↗

Herbaceous Feedstock 2022 State of Technology Report

The U.S. Department of Energy promotes production of advanced liquid transportation fuels from lignocellulosic biomass by funding fundamental and applied research that advances the state of technology (SOT). As part of its involvement in this mission, Idaho National Laboratory completes an annual SOT report for nth-plant and 1st-plant herbaceous biomass feedstock logistics. The purpose of the SOT is to provide the status of feedstock supply system technology development for herbaceous biomass to biofuels relative to technical targets and cost goals from specific design cases, based on data and experimental results. Although conventional feedstock supply systems form the backbone of the emerging biofuels industry, they have limitations that restrict widespread implementation on a national scale. To meet the demands of the future industry, the feedstock supply system must shift from the conventional system to what has been termed “advanced” supply systems. In advanced designs, a distributed network of aggregation and processing centers, termed “depots,” are employed near the points of biomass production (i.e., the field or forest) to reduce feedstock variability and produce feedstocks of a uniform format, moving toward biomass commoditization. The 2022 Herbaceous SOT is part of a vision of achieving an implemented advanced feedstock supply system, which produces a stable, tradable commodity at the decentralized distributed depot. It utilizes feedstock fractionation by incorporating technologies that can separate the biomass into its anatomical fractions (leaves, husks, stems and cobs) to reduce impurities and produce fractions that satisfy downstream quality considerations. By using a series of air classification steps, this strategy can reduce the extrinsic ash in corn stover and produce enriched tissue fractions that can be blended to a conversion specification or converted individually in optimized biochemical conversion campaigns. Additionally, a majority of the leaves (which do not meet the quality specification) are separated out early and can be supplied to alternate markets. The 2022 Herbaceous SOT incorporates an advanced biomass fractionation and processing system to produce pellets enriched tissues from three-pass corn stover. The resulting enriched pellets are delivered to the biorefinery individually where they can be blended to a specification or converted in campaigns where the conditions are optimized for each tissue. Unused fractions can be sent to a a midstream market or to a different conversion process that is better suited to their properties to offset the cost of the delivered feedstock. The main benefits from the proposed system can be summarized as: (1) $6.86/dry ton (2016$) lower cost for the air classification due to elimination of the requirement to discard the high ash lights fraction; (2) $1.56/dry ton lower delivered cost by selling the unsuitable leaf fraction into the feed market as a midstream co-product (assuming a selling price that is 11% higher than their cost of production); (3) 0.98% increase in carbohydrate content (from 60.16% to 61.14%); and (4) 0.97% decrease in ash content (from 6.00% to 5.03%) compared to the 2021 Herbaceous SOT. Overall, the 2022 nth-plant Herbaceous SOT predicts a modeled delivered feedstock cost of $78.64/dry ton (2016$) if it is assumed that the enriched leaf fraction is sold at its production cost; this is a slight increase of $0.43/dry ton increase from the 2021 Herbaceous SOT nth-Supply case cost. The increased cost derived from a $0.38/dry ton increase in transportation and handling cost to procure more biomass (to replace the enriched leaf fraction that was not delivered to the biorefinery. The total preprocessing cost was $0.27/dry ton higher than the 2021 result because of updates to energy consumption, purchasing price and dry matter loss data for the rotary shear ($3.00/dry ton increase) and the pelleting mill ($4.52/dry ton increase). The data utilized were generated in pilot-scale tests in the Biomass Feedstock National User Facility (BFNUF) at INL and at Forest Concepts, including tests for rotary shear and pelleting of the air classified fractions. A greenhouse gas emissions analysis was performed by Argonne National Laboratory using the most up to date version of the Greenhouse Gases, Regulated Emissions, and Energy use in Transportation model (GREET®). The analysis showed an increase of 17.34 kg CO2e/dry ton from the 2021 SOT (67.71 kg CO2e/ton in the 2021 Herbaceous SOT to 85.05 kg CO2e/ton in the 2022 Herbaceous SOT). The net increase is primarily attributed to increased energy consumption in pelleting mill.

09 BIOMASS FUELS↗

Improving climate model coupling through a complete mesh representation: a case study with E3SM (v1) and MOAB (v5.x)

One of the fundamental factors contributing to the spatiotemporal inaccuracy in climate modeling is the mapping of solution field data between different discretizations and numerical grids used in the coupled component models. The typical climate computational workflow involves evaluation and serialization of the remapping weights during the preprocessing step, which is then consumed by the coupled driver infrastructure during simulation to compute field projections. Tools like Earth System Modeling Framework (ESMF) and TempestRemap offer capability to generate conservative remapping weights, while the Model Coupling Toolkit (MCT) that is utilized in many production climate models exposes functionality to make use of the operators to solve the coupled problem. However, such multistep processes present several hurdles in terms of the scientific workflow and impede research productivity. In order to overcome these limitations, we present a fully integrated infrastructure based on the Mesh Oriented datABase (MOAB) library, which allows for a complete description of the numerical grids and solution data used in each submodel. Through a scalable advancing-front intersection algorithm, the supermesh of the source and target grids are computed, which is then used to assemble the high-order, conservative, and monotonicity-preserving remapping weights between discretization specifications. The Fortran-compatible interfaces in MOAB are utilized to directly link the submodels in the Energy Exascale Earth System Model (E3SM) to enable online remapping strategies in order to simplify the coupled workflow process. We demonstrate the superior computational efficiency of the remapping algorithms in comparison with other state-of-the-science tools and present strong scaling results on large-scale machines for computing remapping weights between the spectral element atmosphere and finite volume discretizations on the polygonal ocean grids.

58 GEOSCIENCES↗

Using Temporal Information from Human Mobility Data to Detect Anchor Points

Spatiotemporal mobility data are available in massive quantities, but large quantities of data typically include fewer variables or data fields. Often, the only available fields are User ID, Longitude, Latitude, Timestamp (ULLT). This raises an important question: how much can we infer about human mobility patterns using only these four fields? With ULLT data, we do not know individuals' socioeconomic status information or when they are visiting their anchor points (AP) or locations (such as homes, places of employment, or schools), and it is a modern challenge to use this data to infer these characteristics. When detecting anchor locations with limited input information, verification and validation (VV) are significant challenges. This paper addresses the problem of identifying individuals' anchor locations using only temporal information from spatiotemporal datasets with limited attributes. Our approach does not explicitly use latitude and longitude during analysis. Locationbased information is only employed in the preprocessing stage to identify periods of movement (trips) and stops (dwelling). Beyond this step, all analysis is based on temporal patterns. In theory, if stops and dwell times could be detected through alternative means, our method could function entirely without location-based input. We demonstrate this methodology on the 2017 National Household Travel Survey (NHTS) data, because it includes a carefully designed and collected time use survey with representative sampling and labeled ground truth. The high-quality survey data allows us to test the accuracy of our methods because NHTS contains intended place labels and agent/user characteristics. We have also applied our validated AP identification algorithm on very large-scale GPS based trajectory data for Patterns-of-Life (PoL) assessment and other applications, but due to space limit that could not be presented here.

McBride, Liz [ORNL] (ORCID:0000000286925869)↗

Feedstock Harvesting and Storage: Post-Harvest Management for Quality Preservation

Research on harvesting, collection, and storage has matured, and our understanding of feedstock supply challenges has led us to the point of our current goals. The goals of this project include using storage time to collect impactful data of critical feedstock properties in storage and to characterize the feedstock before delivery. This data will inform delivery scheduling of biomass and relevant compositional information to minimize variations in the material quality, which are critical to the preprocessing and conversion performance.

09 BIOMASS FUELS↗

Forming a database to study reversed magnetic shear from the National Spherical Torus eXperiment using machine learning

Achieving a long-lived reversed magnetic shear (RMS) target plasma in the National Spherical Torus eXperiment Upgrade will require developing various sustainment scenarios. To help with the ongoing plasma control efforts, the development of a new analysis for the motional Stark effect (MSE) diagnostic using a machine learning algorithm, namely, MSE-ML, is described. MSE-ML will be used to identify patterns during RMS discharges, some of which suffer magnetohydrodynamic (MHD) events resulting in current redistribution and monotonic q-profiles. A database consisting of q and magnetic shear profiles is being constructed primarily based on the existing National Spherical Torus eXperiment data with equilibrium reconstructions constrained by the magnetic field pitch angle profile measured using the multi-channel MSE diagnostic. An unsupervised k-means clustering of the data is developed to study the RMS formation as a function of time. The initial clustering from the q-profiles shows significant differences in both amplitude and the duration of the RMS period. As a goal, the clustering results that detect and distinguish shots with substantial and sustained RMS are to be used as a preprocessing step in a supervised algorithm to identify the underlying conditions that lead to long-lasting improved confinement with RMS. Another aim of the MSE-ML study is to identify precursors of RMS-destroying MHD events in either derived data such as the q-profile or directly measured data such as the magnetic field pitch angle profile.

Uzun-Kaymak, I. U. (ORCID:0000000276251493)↗