Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hierarchical data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Phlex: Parallel, Hierarchical, and Layered EXecution of data-processing algorithms

Phlex is a computing framework supporting the parallel, hierarchical, and layered execution of data-processing algorithms. It is based on the functional-programming paradigm, thus guaranteeing thread-safety when invoking user-defined pure functions. Phlex allows users to specify arbitrary graph-based hierarchies of data organization, enabling more flexible processing of data as required by the constraints of the program.

Knoepfel, KyleJ. [Fermi National Accelerator Labor↗

Fox Trails

1. This software utilizes python pandas to pull data from P6 databases or XER files. The software transforms the datasets into multiple main tables by joining, filtering, iteratively flattening hierarchical structured data, and pivoting datasets to give simple flat output tables. The activity table includes all of the information related to an activity including activity codes, global, EPS, and project codes, UDFs, and WBS information as separate columns. This includes the code id, code value and sequence number for all levels in hierarchical codes. The resource table is similar to the activity table and includes all of the information related to resources on activities including UPFs and resource codes. The resource time phased table takes the resource information and time phases it for the budget, forecast, late, and actual dates/units/costs that closely matches P6's user interface's values as it implements the resource curve and calendars. The wbs table contains the WBS structure broken out by levels and includes UDFs, codes, and notebook topics. The final P6 data table is the relationships table which simply contains the relationships. 2. When a user updates the tool with data (via giving it P6 project names with database username/password information or XER files) the system creates the data in #1, then creates a networkx graph with the activity data imbedded in the node data and the relationships added as edges. Each edge also has it's float calculated (working time distance between the predecessor and successor) and attached to the edge. Activities are also tagged as a potential start of a path based on their constraints, constraint dates, remaining start date, and activity status. When a user enters an activity ID into the UI, it runs a shortest path calculation on the network graph between each node tagged as potential start to the entered activity id based on the float tagged on the edge. Each path returned by the algorithm contains all of the nodes on the path in order, as well as the total float of the edges that make the path. This data is then collected and returned to the user in the form of a gantt chart with groupings for each path that includes the total float for each group. 3. Similar to 2, if the user passes through a reference dataset each activity set in the path is checked to see if it had a path in the reference dataset, if that path was the primary path between the start and end activities, and what has changed regarding logic and durations. These changes are color coded and summarized before sent to the user to be displayed by the UI for simple discovery. 4. Utilizing the data from #1, the user can submit desired grouping code(s) and filters to the system. The system will then pull the activities, resources, and relationships and create a gantt chart based on the groupings sent and filtered based on the filters sent. 5. The system will produce a gantt chart in a similar method to #4, but allows interactivity with the data. As the user interacts with the gantt chart, the software captures the changes and stores it with the user making the change so that project controls and implement those changes in P6.

Fox, Ben↗

Self-Supervised and Interpretable Anomaly Detection Using Network Transformers

Machine learning and deep neural networks (DNNs) have been proposed as a tool to identify anomalies in computer network communications. However, due the obfuscated nature of off-the-shelf machine learning models, their output often does not provide enough information to isolate the source of the anomaly to take corrective measures. In this article, we introduce the network transformer (NeT), a DNN model for anomaly detection that incorporates the graph structure of the communication network in order to improve interpretability. Further, the presented approach has the following advantages: first, enhanced interpretability by incorporating the graph structure of computer networks; second, provides a hierarchical set of features that enables analysis at different levels of granularity; second, self-supervised training that does not require labeled data. The NeT model was evaluated on a set of anomalous scenarios executed in a real industrial control system. The presented approach successfully identified the anomalies, the devices affected, and the specific connections causing the anomalies, providing a data-driven hierarchical approach to analyze the behavior of a cyber network.

97 MATHEMATICS AND COMPUTING↗

Dark Energy Survey Year 3 results: optimized $w$CDM simulation-based inference with weak lensing map-level hybrid statistics

We present cosmological constraints from the Dark Energy Survey Year 3 (DES Y3) weak lensing data using hierarchical hybrid statistics within a Bayesian simulation-based inference framework that is based on the Gower Street simulations. To maximize the precision of the inference, we have developed a new, information-theory based, data compression of the weak lensing maps to just seven highly informative summary statistics. The hybrid scheme exploits the high information content of the power spectrum, compressing both the power spectrum and neural-based summaries that are designed to extract further information. Our simulation-based approach enables principled forward modelling of all major sources of systematic uncertainty and survey properties into realistic mock observations, including the survey mask, photometric redshift uncertainties, intrinsic galaxy alignments, multiplicative shear calibration bias, source galaxy clustering, non-Gaussian shape noise, and non-linear structure formation. The summary statistics are then used in a Bayesian simulation-based inference pipeline. The inference is validated through coverage tests and checks for robustness against baryonic feedback. Assuming a $w$CDM cosmology, our analysis yields $S_8 = 0.808 \pm 0.017$, $Ω_{\rm m} = 0.325 \pm 0.024$, and $w < -0.766$ (marginalized posterior 68 per cent credible intervals). This rigorous combination of information theory, physics- and neural network-based extreme data compression, and principled Bayesian analysis improves the figure of merit for $(Ω_{\rm m}, S_8, w)$ by 60 per cent over the previous state-of-the-art, and by almost a factor of 3 over two-point analyses of the same data. They are the most precise joint constraints on $(Ω_{\rm m}, S_8, w)$ from weak gravitational lensing data alone of any survey to date. We intend to apply this analysis to the more recent DES Y6 data.

Williamson, J. [University Coll. London]↗

A seamless approach for evaluating climate models across spatial scales

In regions of the world where topography varies significantly with distance, most global climate models (GCMs) have spatial resolutions that are too coarse to accurately simulate key meteorological variables that are influenced by topography, such as clouds, precipitation, and surface temperatures. One approach to tackle this challenge is to run climate models of sufficiently high resolution in those topographically complex regions such as the North American Regionally Refined Model (NARRM) subset of the Department of Energy’s (DOE) Energy Exascale Earth System Model version 2 (E3SM v2). Although high-resolution simulations are expected to provide unprecedented details of atmospheric processes, running models at such high resolutions remains computationally expensive compared to lower-resolution models such as the E3SM Low Resolution (LR). Moreover, because regionally refined and high-resolution GCMs are relatively new, there are a limited number of observational datasets and frameworks available for evaluating climate models with regionally varying spatial resolutions. As such, we developed a new framework to quantify the added value of high spatial resolution in simulating precipitation over the contiguous United States (CONUS). To determine its viability, we applied the framework to two model simulations and an observational dataset. We first remapped all the data into Hierarchical Equal-Area Iso-Latitude Pixelization (HEALPix) pixels. HEALPix offers several mathematical properties that enable seamless evaluation of climate models across different spatial resolutions including its equal-area and partitioning properties. The remapped HEALPix-based data are used to show how the spatial variability of both observed and simulated precipitation changes with resolution increases. This study provides valuable insights into the requirements for achieving accurate simulations of precipitation patterns over the CONUS. It highlights the importance of allocating sufficient computational resources to run climate models at higher temporal and spatial resolutions to capture spatial patterns effectively. Furthermore, the study demonstrates the effectiveness of the HEALPix framework in evaluating precipitation simulations across different spatial resolutions. This framework offers a viable approach for comparing observed and simulated data when dealing with datasets of varying spatial resolutions. By employing this framework, researchers can extend its usage to other climate variables, datasets, and disciplines that require comparing datasets with different spatial resolutions.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)↗

HDF5 in the exascale era: Delivering efficient and scalable parallel I/O for exascale applications

Accurately modeling real-world systems requires scientific applications at exascale to generate massive amounts of data and manage data storage efficiently. However, parallel input and output (I/O) faces challenges due to new application workflows and the state-of-the-art memory, interconnect, and storage architectures considered in exascale designs. The storage hierarchy has expanded with node-local persistent memory, solid-state storage, and traditional disk and tape-based storage, thus requiring efficiency at each layer and much more efficient data movement among these layers. This paper discusses how the ExaHDF5 project improved the I/O performance and data management for exascale architectures by enhancing HDF5, a widely used parallel I/O library. The team developed an Asynchronous I/O Virtual Object Layer (VOL) connector that allowed overlapping I/O with computation. They also created a Cache VOL to complement asynchronous I/O by incorporating fast storage layers, such as burst buffer and node-local storage, into the parallel I/O workflow through caching and staging data. Additionally, the team enabled data aggregation and I/O at the node level by using a Subfiling Virtual File Driver (VFD). To demonstrate superior I/O performance with HDF5 at exascale, the ExaHDF5 team collaborated with several exascale applications. In this paper, we show I/O performance improvements for three applications: Cabana (a particle-based simulation library), EQSIM (a regional earthquake simulation software), and E3SM (a climate system modeling library).

Asynchronous I/Ol↗

Ecological connectivity and habitat loss shape patterns of genetic diversity in a threatened salamander

Context The maintenance of genetic diversity is essential for preserving adaptive potential in populations, yet it is increasingly threatened by landscape alteration. The field of landscape genetics offers a framework for assessing how patch-level landscape conditions, modeled at multiple scales, influence genetic diversity. Objectives We sought to assess how local environmental features and connectivity influence genetic diversity across 74 four-toed salamander (Hemidactylium scutatum) breeding wetlands in the southeastern United States. Methods Using next-generation sequencing data and hierarchical Bayesian models, we examined genome-wide heterozygosity in relation to local landscape features and ecological connectivity. We also assessed the scale of effect of landscape features and tested for temporal lag effects. Results Genetic diversity was lower in wetlands with higher levels of historic deforestation and lower connectivity. An interaction between deforestation and connectivity indicated that deforestation had stronger negative effects in isolated wetlands but weaker effects in well-connected wetlands. Accounting for scale of effect and temporal lags was critical for detecting these relationships. Conclusions Our analyses highlight the importance of assessing the spatial scale (scale of effect) and temporal lag of landscape features to detect key drivers of genetic diversity. In line with population genetic theory, our results indicate that the genetic consequences of habitat loss do not affect populations uniformly and are most severe in isolated populations where gene flow cannot buffer against loss of diversity. Altogether, we highlight the importance of considering the interaction of habitat loss and connectivity in conservation genetic management.

Hemidactylium scutatum↗

Performance Evaluation of Next-Generation Grid Automation and Controls with High PV Penetration

This paper presents a hardware-in-the-loop (BIL) simulation to evaluate the performance of an advanced grid automation architecture, referred to as data-enhanced hierarchical control (DEHC), in achieving voltage regulation and conservation voltage reduction (CVR) in distribution networks with very high photovoltaic (PV) generation. This architecture comprises an advanced distribution management system (ADMS), a distributed energy resource management system (DERMS), and grid-edge devices working synergistically to provide the grid benefits. The HIL setup used for the evaluation includes ADMS, DERMS, and grid-edge devices. The DEHC performance is evaluated in two representative scenarios considering loose and tight constraints of the power factor at the substation. The results show that the DEHC architecture is effective in achieving voltage regulation and CVR and thus enables the grid integration of high levels of PV generation.

ADMS↗

Scalable Comparative Visualization of Ensembles of Call Graphs

Optimizing the performance of large-scale parallel codes is critical for efficient utilization of computing resources. Code developers often explore various execution parameters, such as hardware configurations, system software choices, and application parameters, and are interested in detecting and understanding bottlenecks in different executions. They often collect hierarchical performance profiles represented as call graphs, which combine performance metrics with their execution contexts. The crucial task of exploring multiple call graphs together is tedious and challenging because of the many structural differences in the execution contexts and significant variability in the collected performance metrics (e.g., execution runtime). In this paper, we present Ensemble CallFlow to support the exploration of ensembles of call graphs using new types of visualizations, analysis, graph operations, and features. We introduce ensemble-Sankey , a new visual design that combines the strengths of resource-flow (Sankey) and box-plot visualization techniques. Whereas the resource-flow visualization can easily and intuitively describe the graphical nature of the call graph, the box plots overlaid on the nodes of Sankey convey the performance variability within the ensemble. Our interactive visual interface provides linked views to help explore ensembles of call graphs, e.g., by facilitating the analysis of structural differences, and identifying similar or distinct call graphs. Finally, we demonstrate the effectiveness and usefulness of our design through case studies on large-scale parallel codes.

97 MATHEMATICS AND COMPUTING↗

Coincident Capture through Post-processing PTRAC [Slides]

This presentation discusses the new PTRAC capabilities and workflows. The PTRAC capability in MCNP6.3 has seen a massive overhaul since MCNP6.2. The new HDF5 file format allows for both MPI- and thread-based parallelism. MCNPTools has been updated to handle the new HDF5 PTRAC format and is now open sourced on GitHub. Built-in capabilities, such as the pulse-height tally coincident capture special treatment, can largely be replicated through separate postprocessing scripts that leverage both PTRAC and MCNPTools. This allows for greater flexibility in user-specified and controlled detector response functionality, ultimately using the MCNP code for what it is best at (i.e., particle transport).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Simultaneous inference of equation of state parameters and unknown data errors with uncertainty quantification via hierarchical Bayesian posterior maximization

Equations of state (EOSs) are a key component in running hydrodynamic simulations as they relate the thermodynamic states for the material. The Davis reactants EOS is commonly used for modeling high explosives (HEs), and the EOS model parameters are calibrated using material specific data. The calibrations are often performed with uncertainty quantification via Bayesian inference to account for uncertainty in the data and generate ensembles of likely parameters. However, there are relatively few HE data sets to use for calibration and many are historical and lack error information. In this work, we simultaneously calibrate the Davis reactants EOS model parameters and unknown data error terms for the high explosive PBX 9501. To quantify the uncertainty in the models and the data, we use a Bayesian framework for the calibration and compute the hierarchical Bayesian posterior distribution with both a posteriori maximization approach and Markov Chain Monte Carlo. In general, we find that, given our assumptions, the two approaches result in similar calibrated parameters, posterior covariance matrices, and insights about the parameters but that the posterior maximization requires far less computational resources.

97 MATHEMATICS AND COMPUTING↗

PaleoSTeHM v1.0: a modern, scalable spatiotemporal hierarchical modeling framework for paleo-environmental data

Abstract. Geological records of past environmental change provide crucial insights into long-term climate variability, trends, non-stationarity, and nonlinear feedback mechanisms. However, reconstructing spatiotemporal fields from these records is statistically challenging due to their sparse, indirect, and noisy nature. Here, we present PaleoSTeHM, a scalable and modern framework for spatiotemporal hierarchical modeling of paleo-environmental data. This framework enables the implementation of flexible statistical models that rigorously quantify spatial and temporal variability from geological data while clearly distinguishing measurement and inferential uncertainty from process variability. We illustrate its application by reconstructing temporal and spatiotemporal paleo-sea-level changes across multiple locations. Using various modeling and analysis choices, PaleoSTeHM demonstrates the impact of different methods on inference results and computational efficiency. Our results highlight the critical role of model selection in addressing specific paleo-environmental questions, showcasing the PaleoSTeHM framework's potential to enhance the robustness and transparency of paleo-environmental reconstructions.

58 GEOSCIENCES↗

Hierarchical Bayesian modeling for Inverse Uncertainty Quantification of system thermal-hydraulics code using critical flow experimental data

The best estimate plus uncertainty methodology in nuclear system thermal-hydraulic studies necessitates a comprehensive understanding of uncertainties in system code predictions. The forward uncertainty quantification (UQ) process involves the propagation of input uncertainties through the computational models to obtain uncertainties in the outputs. To this end, achieving an accurate estimation of input uncertainties is important, which is the focus of inverse UQ (IUQ). Traditionally, research in Bayesian IUQ within the nuclear engineering domain has largely relied on single-level Bayesian inference. While being effective for relatively small datasets, this approach encounters limitations for cases with large datasets. The use of a single-level model may prove inefficient, as the resultant posterior distributions can significantly differ when distinct subsets of data are employed. To address this issue, we employ an hierarchical Bayesian model for IUQ. Furthermore, this approach involves organizing observations into different groups based on the test conditions, thereby accommodating varying calibration parameters across these distinct groups. In this study, we developed and implemented a hierarchical Bayesian IUQ method to consider the grouping effect of critical flow measurement data from various geometries. Comparing the outcomes of IUQ under different selections of test data using hierarchical Bayesian IUQ against those obtained from single-level Bayesian IUQ, the forward propagation of hierarchical Bayesian IUQ results demonstrates a notably improved agreement with the experimental data.

42 ENGINEERING↗

Discrepancy quantification between experimental and simulated data of CO 2 adsorption isotherm using hierarchical Bayesian estimation

Here, to quantitatively analyze the inconsistencies commonly observed between experimental and simulated adsorption isotherms, parameter estimation of adsorption isotherm models was conducted by hierarchical Bayesian estimation with parameter uncertainties being quantified as probability distributions. The estimation method was implemented using Markov Chain Monte Carlo (MCMC) to analyze multiple data sets obtained from different sources, including a publicly available database. To describe the discrepancies of experimental and simulated adsorption data, the simulation data was set as the reference to which experimental measurements were compared. We applied the proposed approach to analyze CO 2 adsorption isotherms that are measured and simulated on zeolite 13X and MIL-101(Cr). In these case studies, the discrepancy of CO 2 adsorption isotherm was successfully quantified between experimental measurements and predictions given by molecular simulations using Grand Canonical Monte Carlo (GCMC), where uncertainties were quantified as probability distributions. Furthermore, experimental data sets that agree well with the GCMC simulation have been identified, providing insights into experimental and measurement methods as well as choosing the right assumptions in the molecular simulation.

42 ENGINEERING↗

DeepLynx Ecosystem 2025

Poor data integration and governance continue to plague complex engineering projects, resulting in missed cost, schedule, and performance targets. Departments operate in isolated systems with manual data exchange, creating fragmented information that compounds errors and leads to significant delays and cost overruns. The DeepLynx ecosystem addresses these challenges through an open-source, modular data management platform that transforms fragmented project data into an integrated digital thread. Built on a federated microservice architecture, the ecosystem comprises seven specialized tools centered around DeepLynx Nexus, a unified data catalog with hierarchical organization and graph-based navigation capabilities. The ecosystem includes: DeepLynx Stream for real-time timeseries data ingestion from industrial sources; DeepLynx Ingest for governed data uploads with formal review workflows; DeepLynx Lattice for ontology-based entity and relationship extraction; DeepLynx Run for workflow orchestration and secure AI/ML compute; DeepLynx Visualize for 3D digital twin visualization; and DeepLynx Insight for AI-assisted document analysis with traceable, grounded responses. Deployable in cloud, on-premise, or hybrid environments using containerized Docker applications and Helm charts, the DeepLynx ecosystem provides flexible infrastructure that adapts to organizational requirements. By consolidating project data into a unified data lake with role-based access controls and OAuth2 authentication, DeepLynx enables digital thread and digital twin capabilities that improve decision-making, reduce risk, and support complex engineering workflows throughout the project lifecycle.

42 - ENGINEERING↗