Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data visualizations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Ionospheric Disturbances in GNSS TEC Data: SpaceX Falcon 9 Deorbit Maneuvers Over CONUS in April–May 2024

Traveling ionospheric disturbances (TIDs) driven by a large number of internal and external sources are detectable with dense networks of ground‐based Global Navigation Satellite System (GNSS) receivers' measurements of total electron content (TEC). We present the newly developed System for Rapid Analysis of Ionospheric Dynamics (S‐RAID), providing data and visual GNSS TEC products of TIDs of periods ~2–120 min and horizontal resolution up to tens of kilometers for years 2017–2024. The S‐RAID data reveal myriad natural and anthropogenic TIDs from meteorology and space weather, and from human spaceflight activities. In this report, we focus on new prominent disturbances found during SpaceX second stage deorbit maneuvers in April–May 2024. Signatures include waves emanating from the Falcon's trajectory above California and depletions following its deorbit and passage over Arizona. These findings suggest further opportunities to detect and quantify small‐scale events in the ionosphere as well as to understand the responses of the atmosphere‐ionosphere system to known inputs.

58 GEOSCIENCES↗

The Exascale Framework for High Fidelity coupled Simulations (EFFIS): Enabling whole device modeling in fusion science

We present the Exascale Framework for High Fidelity coupled Simulations (EFFIS), a workflow and code coupling framework developed as part of the Whole Device Modeling Application (WDMApp) in the Exascale Computing Project. EFFIS consists of a library, command line utilities, and a collection of run-time daemons. Together, these software products enable users to easily compose and execute workflows that include: strong or weak coupling, in situ (or offline) analysis/visualization/monitoring, command-and-control actions, remote dashboard integration, and more. We describe WDMApp physics coupling cases and computer science requirements that motivate the design of the EFFIS framework. Furthermore, we explain the essential enabling technology that EFFIS leverages: ADIOS for performant data movement, PerfStubs/TAU for performance monitoring, and an advanced COUPLER for transforming coupling data from its native format to the representation needed by another application. Finally, we demonstrate EFFIS using coupled multi-simulation WDMApp workflows and exemplify how the framework supports the project’s needs. We show that EFFIS and its associated services for data movement, visualization, and performance collection does not introduce appreciable overhead to the WDMApp workflow and that the resource-dominant application’s idle time while waiting for data is minimal.

97 MATHEMATICS AND COMPUTING↗

TEACHING AN OLD ACCELERATOR NEW TRICKS

The Argonne Tandem Linac Accelerator System (ATLAS) has been a National User Facility since 1985. In that time, many of the systems that help operators retrieve, modify, and store beamline parameters have not kept pace with the advancement of technology. Development of a new method of storing and retrieving beamline parameters resulted in the testing and installation of a time-series database as a potential replacement for the traditional relational database. InfluxDB was selected due to its self-hosted Open-Source version availability as well as the simplicity of installation and setup. A program was written to periodically gather all accelerator parameters in the control system and store them in the time-series database. This resulted in over 13,000 distinct data points, captured at 5-minute intervals. A second test captured 35 channels on a 1-minute cadence. Graphing of the captured data is being done on Grafana, an Open-Source version is available that co-exists well with InfluxDB as the back-end. Grafana made visualizing the data simple and flexible. The testing has allowed for the use of modern graphing tools to generate new insights into operating the accelerator, as well as opened the door to building large data sets suitable for Artificial Intelligence and Machine Learning applications.

Novak, D.↗

Towards Autonomous Experiments by Connecting High Performance Microscopy with High Performance Computing

The digitization of controls, data, and analysis in microscopy is bringing the idea of autonomous microscopes closer to reality than ever before. Automated transmission electron microscopy (TEM) is already fairly routine for some experiments the only require simple repetitive tasks such as imaging biological macromolecules for single particle cryoEM [1], tilt series for electron tomography [2], and movies for crystallography [3]. The vast majority of TEM experiments are conducted completely by human operators who choose the regions of interest, optimize experimental parameters, and make decisions about data quality visually during an experiment. The field is still a long way from having completely autonomous TEMs that can adapt to sample difficulties and tune experimental parameters based on data quality and desired experimental outcomes. Part of the issue is the lack of capability for feeding information learned from on-line, live data analysis back into the on-going experiment [4]. Furthermore, this presentation will discuss current capabilities for large scale data reduction and analysis using high performance computing (i.e. supercomputing) and progress towards developing a true feed-back loop that places data analysis and theory in the experimental loop.

97 MATHEMATICS AND COMPUTING↗

Web-Based Tools for Data-Informed Remedy Optimization: Software Theory and User Guide

This report documents the development and application of two web-based decision-support tools for pump-and-treat (P&T) groundwater remediation systems: PTOLEMY (Pump-and-Treat Optimized Location Evaluation to Maximize Yields) and OPTIMA (Optimization for Pump-and-Treat Implementation, Management, & Assessment). These tools enhance remedy design and management by leveraging advanced computational methods – specifically deep learning and multi-objective optimization – within a user-friendly platform. By integrating data-driven models with established hydrogeological knowledge, PTOLEMY and OPTIMA enable more efficient evaluation of well placement and operational strategies, helping site managers balance multiple remediation objectives under complex conditions. Both tools are implemented as modules within the SOCRATES (Suite Of Comprehensive Rapid Analysis Tools for Environmental Sites) web platform, which provides data access, visualization, and analytics to support remedy optimization across sites in the U.S. Department of Energy Office of Environmental Management complex. PTOLEMY is a rapid screening module designed to identify promising locations for new extraction wells. It employs a multi-channel three-dimensional convolutional neural network (MC3D-CNN) trained on high-fidelity simulation data to predict the relative performance (in terms of contaminant mass recovery) of potential well sites. Through an interactive web interface, PTOLEMY visualizes the probability of high performance across a site, highlighting areas where an extraction well is likely to yield above-threshold contaminant removal over a multi-year period. PTOLEMY’s map-based displays and exportable results support transparent communication of screening analyses. By focusing attention on the most favorable candidate locations, the tool augments traditional engineering judgment and physics-based modeling, providing a data informed basis for subsequent detailed evaluations. OPTIMA is a multi objective optimization module designed to find wellfield layouts and operating schedules that meet various cleanup goals. It quickly evaluates thousands of candidate setups – combinations of well locations, timing, and rates – and returns a small set of best trade-off options for comparison. At its core, OPTIMA uses a U-Net-based surrogate model – a deep-learning emulator of a groundwater flow and transport simulator – to dramatically accelerate scenario evaluations. Coupling this fast surrogate with the NSGA-II (Non-dominated Sorting Genetic Algorithm II) evolutionary algorithm, OPTIMA explores a wide decision space of well locations and schedules to identify Pareto-optimal solutions that trade off key objectives (e.g., minimizing cleanup time, maximizing contaminant mass removal, and minimizing plume extent). The tool outputs a family of optimal configurations and visualizes their trade-offs (Pareto frontiers of cleanup metrics and maps of optimized well placements). Site managers can use these results to understand the range of viable strategies and to select candidate designs for more detailed verification. OPTIMA is currently under active development and not yet fully released; this guide provides early documentation to support planning and gather user feedback.

54 ENVIRONMENTAL SCIENCES↗

Coal Power Plant Reinvestment Visualization Tool

This tool, available at https://energycommunities.gov/coal-power-plant-reinvestment-visualization-tool/, serves as a public database and map for the purposes of enabling state and local economic development officials, project developers, and power plant owners to identify and pursue opportunities for plant and community reinvestment. The Coal Power Plant Reinvestment Visualization Tool focuses on coal power plants that have been closed or set-to-retire, alongside key infrastructure characteristics that are relevant for potential redevelopment reinvestment opportunities, including the opportunity to query these data based off pre-defined or user-set queries to identify opportunities for coal power facility reinvestment to support solar, wind, manufacturing, and nuclear. These Data for visualization and query include, but are not limited to: • Electric Transmission Lines • EPA Brownfield Sites (assessed with Federal funding) • Petroleum Terminals • Ports • Railroads • Others The tool will be updated periodically to provide relevant information to enable state and local economic development officials, project developers, and power plant owners to identify and pursue opportunities for economic revitalization and community reinvestment. Within this application data are focused on U.S. power plants with coal generation, where at least one on-site generator utilizes or utilized coal as a fuel source, including planned retirements through 2058 (EIA, 2024). Additional information on energy communities and revitalization opportunities can be accessed on the Interagency Working Group on Coal & Power Plant Communities & Economic Revitalization Energy Communities website (https://energycommunities.gov).

Closures↗

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES↗

HERO WEC 2024 - Electrical Configuration Deployment Data

The following submission includes raw and processed electrical configuration deployment data from the in water deployment of NREL's Hydraulic and Electric Reverse Osmosis Wave Energy Converter (HERO WEC), in the form of parquet files, TDMS files, CSV files, bag files, and MATLAB workspaces. This dataset was collected in April 2024 at the Jennette's pier test site in North Carolina. Raw data as TDMS, CSV, and bag files are provided here alongside processed data in the form of MATLAB workspaces and Parquet files. This dataset includes the Python code used to process the data and MATLAB scripts to visualize the processed data. All data types, calculations, and processing is described in the included "Data Descriptions" document. All files in this dataset are described in detail in the included README. This data set has been developed by the National Renewable Energy Laboratory, operated by Alliance for Sustainable Energy, LLC, for the U.S. Department of Energy (DOE) under Contract No. DE-AC36-08GO28308. Funding provided by the U.S. Department of Energy Office of Energy Efficiency and Renewable Energy Water Power Technologies Office.

16 TIDAL AND WAVE POWER↗

Streaming Data in HPC Workflows Using ADIOS

The “IO Wall” problem, in which the gap between computation rate and data access rate grows continuously, poses significant problems to scientific workflows which have traditionally relied upon using the filesystem for intermediate storage between workflow stages. One way to avoid this problem in scientific workflows is to stream data directly from producers to consumers and avoiding storage entirely. However, the manner in which this is accomplished is key to both performance and usability. This paper presents the Sustainable Staging Transport, an approach which allows direct streaming between traditional file writers and readers with few application changes. SST is an ADIOS “engine”, accessible via standard ADIOS APIs, and because ADIOS allows engines to be chosen at run-time, many existing file-oriented ADIOS workflows can utilize SST for direct application-to-application communication without any source code changes. This paper describes the design of SST and presents performance results from various applications that use SST, for feeding model training with simulation data with substantially higher bandwidth than the theoretical limits of Frontier’s file system, for strong coupling of separately developed applications for multiphysics multiscale simulation, or for in situ analysis and visualization of data to complete all data processing shortly after the simulation finishes.

Podhorszki, Norbert [ORNL] (ORCID:000000019647542X↗

Open Source Platform for Live Processing of High Data Rate Electron Microscopy (SBIR Phase I Award Topic 15, Subtopic B, Final Scientific / Technical Report)

4D scanning transmission electron microscopy (STEM) is experiencing a revolution both in terms of data size, and data veracity, as are many other areas of experimental data acquisition. This presents enormous opportunities to make new scientific discoveries, but is also hamstrung by aging software that is no longer fit for purpose. Experimental data acquisition platforms must make the same investments in software infrastructure that the simulation community made decades ago to take advantage of high-performance computing (HPC) if they hope to realize the full potential of the latest advances made in detector technology. This requires rethinking the design and deployment of hardware, moving away from small, fixed commodity computing capabilities with closed interfaces to open, modular, scalable systems designed for parallelism and throughput. The technical approach will leverage best-in-class open source technologies, and develop new approaches strategically to build an open source platform for experimental data acquisition targeting high data rate electron microscopy. The project will develop a software platform that uses the power and performance of modern HPC resources to ingest the large volume of data produced by electron microscopes using direct electron detectors. A set of configurable pipelines will be used to enable the processing to be tailored to the problem being investigated with builtin visualization of data at interactive rates. This project will leverage investments made by the DOE and other agencies to provide a powerful software platform for the experimental data community working with most powerful electron microscopes, such as 4D STEM detectors. It will create a platform that serves as a reference for the community. Permissive open source licensing will make the platform available to all, with a simple deployment strategy offering the ability to keep data close to where it is generated. This advanced software platform will require customization for each deployment, creating opportunities for software services that are typically offered to extend open source software platforms.

97 MATHEMATICS AND COMPUTING↗

Community Resilience Indicator Analysis: Commonly Used Indicators from Peer-Reviewed Research (Updated for Research Published 2003-2021)

In 2017, FEMA’s National Integration Center (NIC) Technical Assistance (TA) Branch identified a need to establish a data-driven basis for prioritizing locations for TA investment and guiding local emergency management planning. To achieve this goal, FEMA tasked Argonne National Laboratory (Argonne) with identifying commonly used indicators of community resilience across the landscape of published peer-reviewed research. FEMA and Argonne completed the first Community Resilience Indicator Analysis (CRIA) in 2018 and repeated the process in 2022. The CRIA process begins with a literature review and cataloguing of published peer-reviewed assessment methodologies on social vulnerability and community resilience. The literature review findings are then filtered by inclusion criteria established by the CRIA research team to ensure the methodologies are: (1) Quantitative, (2) Data and methodology are publicly available, (3) Calculated at the county level or lower, (4) Examine generalized hazard risk (rather than a singular hazard), and (5) Focused on pre-disaster community conditions. After this, the research team identifies the commonly used indicators across these methodologies and selects the best data source for each indicator. Finally, the research team bins the data for visual display, conducts a correlation analysis and creates a composite index, the FEMA Community Resilience Index (FEMA CRI). In 2018, the CRIA identified eight resilience and vulnerability assessment methodologies and 20 commonly used indicators (indicators used in three or more of the eight methodologies). The FEMA CRI in 2018 was created from these 20 indicators and was produced for at the county level. The 2022 CRIA updated the literature review to expand the list of methodologies examined and followed the same process, resulting in an analysis of 14 methodologies published between 2003 and 2021 and 22 indicators identified as commonly used (indicators used in five or more of the 14 methodologies). In 2022, the research team produced the FEMA CRI at the county and the census tract levels. To make the CRIA data more accessible and more actionable, each individual indicator and the FEMA CRI is binned and included in FEMA’s Resilience Analysis and Planning Tool (RAPT). RAPT enables emergency managers and community partners to quickly visualize relative differences in potential resilience by county, tribe and census tract. By reviewing the data for each of these 22 indicators individually, emergency managers can gain insights for targeted outreach strategies, planning, mitigation investments and response and recovery operations. Communities, regional governments and others can use this data to better understand potential challenges to resilience. As the social science field of examining and validating indicators of resilience evolves, FEMA will update RAPT to provide emergency managers and community partners with additional data and tools to inform planning, mitigation, response and recovery. It is important to understand that the role of the emergency manager is not to change or to “improve” the data, but to plan appropriately for the community characteristics reflected in the data. These datasets are community characteristics that researchers have identified as important considerations for resilience. For example, people with disabilities may have greater challenges to be resilient to disasters. If a community has a high population of people with disabilities, the emergency manager(s) may need to create tailored preparedness outreach programs and strategies to ensure those residents have support if evacuation is necessary. Rather than label these indicators as an absolute measure of resilience, FEMA considers “potential challenges to resilience” a better frame to understand these indicators. Everyone is vulnerable to disasters. While scholars theorize that certain characteristics may make an individual or a household more socially vulnerable, the data does not reflect measures that individuals and/or communities have taken to address potential challenges, such as emergency management planning and outreach or household preparedness measures. To aid emergency managers in understanding how to use these indicators, calling them potential challenges to resilience supports a more positive and strategic application of the data in all phases of emergency management.

99 GENERAL AND MISCELLANEOUS↗

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

ALPINE Overview

Data-driven sampling enables probabilistic identification of interesting regions in the data automatically, prioritizing important regions. Applied in situ to Nyx, important halo regions are preserved.

97 MATHEMATICS AND COMPUTING↗

A Graphical Model for Fusing Diverse Microbiome Data

This paper develops a Bayesian graphical model for fusing disparate types of count data. The motivating application is the study of bacterial communities from diverse high-dimensional features, in this case, transcripts, collected from different treatments. In such datasets, there are no explicit correspondences between the communities and each corresponds to different factors, making data fusion challenging. We introduce a flexible multinomial-Gaussian generative model for jointly modeling such count data. This latent variable model jointly characterizes the observed data through a common multivariate Gaussian latent space that parameterizes the set of multinomial probabilities of the transcriptome counts. The covariance matrix of the latent variables induces a covariance matrix of co-dependencies between all the transcripts, effectively fusing multiple data sources. We present a computationally scalable variational Expectation-Maximization (EM) algorithm for inferring the latent variables and the parameters of the model. Here, the inferred latent variables provide a common dimensionality reduction for visualizing the data and the inferred parameters provide a predictive posterior distribution. In addition to simulation studies that demonstrate the variational EM procedure, we apply our model to a bacterial microbiome dataset.

59 BASIC BIOLOGICAL SCIENCES↗

Processed data from the Building Management System for the System Engineering Building.

The dataset spans November 2018 to May 2020 and includes time-series measurements corresponding to supply and return temperatures of air and water, air, hot water and cold water flow rates, energy and power consumption, set-points etc. as a single CSV file. In addition to the measurements, a metadata .json file, and a .ttl file to visualize the data as per BRICK schema are also included.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Vanderbilt Alumni Hall

This dataset includes processed and raw data from the Vanderbilt Alumni Hall Building (VAH) located on the University of Vanderbilt Campus in Nashville, TN. VAH is a mixed use commercial building (LEED Gold Certified) that consists of classrooms, meeting rooms, office rooms, conference rooms, an exercise room, and others. This dataset spans one year of time-series data from 2019 and includes measurements corresponding to whole building electrical power, thermal power, zone level thermal power, set-point, indoor temperature, humidity, water and air side supply and return temperatures and flow rates as CSV files. In addition to the measurements, a metadata .json file, and a .ttl file to visualize the data as per BRICK schema are also included.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

RCSB Protein Data Bank: improved annotation, search and visualization of membrane protein structures archived in the PDB

Abstract Motivation Membrane proteins are encoded by approximately one fifth of human genes but account for more than half of all US FDA approved drug targets. Thanks to new technological advances, the number of membrane proteins archived in the PDB is growing rapidly. However, automatic identification of membrane proteins or inference of membrane location is not a trivial task. Results We present recent improvements to the RCSB Protein Data Bank web portal (RCSB PDB, rcsb.org) that provide a wealth of new membrane protein annotations integrated from four external resources: OPM, PDBTM, MemProtMD and mpstruc. We have substantially enhanced the presentation of data on membrane proteins. The number of membrane proteins with annotations available on rcsb.org was increased by ∼80%. Users can search for these annotations, explore corresponding tree hierarchies, display membrane segments at the 1D amino acid sequence level, and visualize the predicted location of the membrane layer in 3D. Availability and implementation Annotations, search, tree data and visualization are available at our rcsb.org web portal. Membrane visualization is supported by the open-source Mol* viewer (molstar.org and github.com/molstar/molstar). Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗