Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data visualizations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

WELLBASE - An Interactive Platform for Wellbore Material Assessment

This project seeks to build an open-source wellbore material data repository with adequate material performance and contextual data to support Geological Carbon Storage (GCS). By appropriately evaluating the data types as mentioned earlier made available by the WELLBASE tool, stakeholders can make more informed decisions regarding well selections, risk assessment, and economic analysis for geologic carbon storage projects. Advanced Natural Language Processing models and other custom python scripts will be deployed in an automated process to extract unstructured data from documents, reports, and web applications and subsequently parse to more usable formats. The processed data will then be integrated into a robust and comprehensive database architecture, optimizing data accessibility, and usability for analytical purposes. The final data products will be accessible through a user-friendly visualization platform that will allow users to query and visualize the data, as well as download data in usable formats.

Tetteh, Daniel A.↗

Enhancing Data Quality Monitoring at CMS with Interactive Visualization Tools and Automated Reference Run Selection

Current data quality monitoring (DQM) tools at CMS offer granularity limited to per-run analysis. Consequently, issues manifesting at the per-lumisection level can go unnoticed or, even if detectable, often lead to the classification of the whole run as bad, resulting in unnecessary data loss. Additionally, shifters have to evaluate a large set of monitoring elements during their long shifts, increasing the probability of human errors or overlooked problems. In this contribution, we present ongoing work on the development of tools that will provide shifters with an accessible, granularity-enhanced view of DQM data through interactive and dynamic visualizations. Furthermore, we introduce a reference run selection tool currently under development, which will automate the selection based on data-taking conditions and will offer a curated set of training data for machine learning models that will be used for the partial automation of the offline data certification process. These endeavors will be integrated into the DIALS website, enabling enhancements in data certification accuracy and improving the accessibility of DQM at CMS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

NREL’s Cyber-Energy Emulation Platform for Research and System Visualization

This paper presents NREL’s Cyber Energy Emulation (CEE) Platform. This is a novel emulation platform for achieving real-time visualization of large-scale environments involving cyber-physical devices. It allows for the environment to include real, physical hardware, along with emulated devices communicating with each other as part of the same system. The CEE Platform is also capable of streaming, collecting, storing, transporting, and visualizing all data within the emulated environment. By providing this capability, it enables high-fidelity visual analysis of events to be performed in real time as well as the use of historical data for forensic analysis. This paper presents the design of the CEE Platform as well as several potential use cases and applications. It also aims to highlight why this type of visualization tool has potential for research and education about cyber-physical systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Improving Performance of M-to-N Processing and Data Redistribution in In Transit Analysis and Visualization

In an in transit setting, a parallel data producer, such as a numerical simulation, runs on one set of ranks M, while a data consumer, such as a parallel visualization application, runs on a different set of ranks N. One of the central challenges in this in transit setting is to determine the mapping of data from the set of M producer ranks to the set of N consumer ranks. This is a challenging problem for several reasons, such as the producer and consumer codes potentially having different scaling characteristics and different data models. The resulting mapping from M to N ranks can have a significant impact on aggregate application performance. In this work, we present an approach for performing this M-to-N mapping in a way that has broad applicability across a diversity of data producer and consumer applications. We evaluate its design and performance with a study that runs at high concurrency on a modern HPC platform. By leveraging design characteristics, which facilitate an “intelligent” mapping from M-to-N, we observe significant performance gains are possible in terms of several different metrics, including time-to-solution and amount of data moved.

Loring, Burlen↗

Improving Performance of M-to-N Processing and Data Redistribution in In Transit Analysis and Visualization

In an in transit setting, a parallel data producer, such as a numerical simulation, runs on one set of ranks M, while a data consumer, such as a parallel visualization application, runs on a different set of ranks N: One of the central challenges in this in transit setting is to determine the mapping of data from the set of M producer ranks to the set of N consumer ranks. This is a challenging problem for several reasons, such as the producer and consumer codes potentially having different scaling characteristics and different data models. The resulting mapping from M to N ranks can have a significant impact on aggregate application performance. In this work, we present an approach for performing this M-to-N mapping in a way that has broad applicability across a diversity of data producer and consumer applications. We evaluate its design and performance with a study that runs at high concurrency on a modern HPC platform. By leveraging design characteristics, which facilitate an ''intelligent'' mapping from M-to-N, we observe significant performance gains are possible in terms of several different metrics, including time-to-solution and amount of data moved.

Loring, Burlen↗

A Comparative Study of the Perceptual Sensitivity of Topological Visualizations to Feature Variations

Color maps are a commonly used visualization technique in which data are mapped to optical properties, e.g., color or opacity. Color maps, however, do not explicitly convey structures (e.g., positions and scale of features) within data. Topology-based visualizations reveal and explicitly communicate structures underlying data. Although our understanding of what types of features are captured by topological visualizations is good, our understanding of people's perception of those features is not. Further, this paper evaluates the sensitivity of topology-based isocontour, Reeb graph, and persistence diagram visualizations compared to a reference color map visualization for synthetically generated scalar fields on 2-manifold triangular meshes embedded in 3D. In particular, we built and ran a human-subject study that evaluated the perception of data features characterized by Gaussian signals and measured how effectively each visualization technique portrays variations of data features arising from the position and amplitude variation of a mixture of Gaussians. For positional feature variations, the results showed that only the Reeb graph visualization had high sensitivity. For amplitude feature variations, persistence diagrams and color maps demonstrated the highest sensitivity, whereas isocontours showed only weak sensitivity. These results take an important step toward understanding which topology-based tools are best for various data and task scenarios and their effectiveness in conveying topological variations as compared to conventional color mapping.

97 MATHEMATICS AND COMPUTING↗

Quantitatively Monitoring Bubble-Flow at a Seep Site Offshore Oregon: Field Trials and Methodological Advances for Parallel Optical and Hydroacoustical Measurements

Two lander-based devices, the Bubble-Box and GasQuant-II, were used to investigate the spatial and temporal variability and total gas flow rates of a seep area offshore Oregon, United States. The Bubble-Box is a stereo camera–equipped lander that records bubbles inside a rising corridor with 80 Hz, allowing for automated image analyses of bubble size distributions and rising speeds. GasQuant is a hydroacoustic lander using a horizontally oriented multibeam swath to record the backscatter intensity of bubble streams passing the swath plain. The experimental set up at the Astoria Canyon site at a water depth of about 500 m aimed at calibrating the hydroacoustic GasQuant data with the visual Bubble-Box data for a spatial and temporal flow rate quantification of the site. For about 90 h in total, both systems were deployed simultaneously and pressure and temperature data were recorded using a CTD as well. Detailed image analyses show a Gaussian-like bubble size distribution of bubbles with a radius of 0.6–6 mm (mean 2.5 mm, std. dev. 0.25 mm); this is very similar to other measurements reported in the literature. Rising speeds ranged from 15 to 37 cm/s between 1- and 5-mm bubble sizes and are thus, in parts, slightly faster than reported elsewhere. Bubble sizes and calculated flow rates are rather constant over time at the two monitored bubble streams. Flow rates of these individual bubble streams are in the range of 544–1,278 mm 3 /s. One Bubble-Box data set was used to calibrate the acoustic backscatter response of the GasQuant data, enabling us to calculate a flow rate of the ensonified seep area (~1,700 m 2 ) that ranged from 4.98 to 8.33 L/min (5.38 × 10 6 to 9.01 × 10 6 CH 4 mol/year). Such flow rates are common for seep areas of similar size, and as such, this location is classified as a normally active seep area. For deriving these acoustically based flow rates, the detailed data pre-processing considered echogram gridding methods of the swath data and bubble responses at the respective water depth. The described method uses the inverse gas flow quantification approach and gives an in-depth example of the benefits of using acoustic and optical methods in tandem.

54 ENVIRONMENTAL SCIENCES↗

System and method of managing large data files

Disclosed are systems and software that provide a high-performance, extensible file format and web API for remote data access and a visual interface for data viewing, query, and analysis. The described system can support storage of raw spectroscopic data such as neural recording data, MSI data, metadata, and derived analyses in a single, self-describing format that may be compatible by a large range of analysis software.

Bowen, Benjamin P.↗

Information-Theoretic Exploration of Multivariate Time-Varying Image Databases

Modern scientific simulations produce very large datasets, making interactive exploration of such data computationally prohibitive. An increasingly common data reduction technique is to store visualizations and other data extracts in a database. The Cinema project is one such approach, storing visualizations in an image database for post hoc exploration and interactive image-based analysis. This work focuses on developing efficient algorithms that can quantify various types of multivariate dependencies existing within multi-variable datasets. It applies specific mutual information measures for the quantification of salient regions from multivariate image data. Here, using such information measures, the opacity of the images is modulated so that the salient regions are automatically highlighted and the domain scientists can interactively explore the most relevant regions for scientific discovery.

97 MATHEMATICS AND COMPUTING↗

Tractometry of the Human Connectome Project: resources and insights

The Human Connectome Project (HCP) has become a keystone dataset in human neuroscience, with a plethora of important applications in advancing brain imaging methods and an understanding of the human brain. We focused on tractometry of HCP diffusion-weighted MRI (dMRI) data. We used an open-source software library (pyAFQ; https://yeatmanlab.github.io/pyAFQ) to perform probabilistic tractography and delineate the major white matter pathways in the HCP subjects that have a complete dMRI acquisition (n = 1,041). We used diffusion kurtosis imaging (DKI) to model white matter microstructure in each voxel of the white matter, and extracted tract profiles of DKI-derived tissue properties along the length of the tracts. We explored the empirical properties of the data: first, we assessed the heritability of DKI tissue properties using the known genetic linkage of the large number of twin pairs sampled in HCP. Second, we tested the ability of tractometry to serve as the basis for predictive models of individual characteristics (e.g., age, crystallized/fluid intelligence, reading ability, etc.), compared to local connectome features. To facilitate the exploration of the dataset we created a new web-based visualization tool and use this tool to visualize the data in the HCP tractometry dataset. Finally, we used the HCP dataset as a test-bed for a new technological innovation: the TRX file-format for representation of dMRI-based streamlines. We released the processing outputs and tract profiles as a publicly available data resource through the AWS Open Data program's Open Neurodata repository. We found heritability as high as 0.9 for DKI-based metrics in some brain pathways. We also found that tractometry extracts as much useful information about individual differences as the local connectome method. We released a new web-based visualization tool for tractometry—“Tractoscope” (https://nrdg.github.io/tractoscope). We found that the TRX files require considerably less disk space-a crucial attribute for large datasets like HCP. In addition, TRX incorporates a specification for grouping streamlines, further simplifying tractometry analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Utah FORGE: Well 16A(78)-32 Perforation Images and Raw Data

This archive contains raw data of visual and acoustic mapping of perforations in Utah FORGE well 16A(78)-32 acquired during the August 2024 circulation program. The dataset includes downhole images captured by EV, a downhole visual analytics company, providing visual records of each perforation. Images are organized in two folders: one set with perforation visualization overlays and one without. An included Excel spreadsheet provides the organized raw data.

15 GEOTHERMAL ENERGY↗

UAE6 - Wind Tunnel Tests Data - UAE6 - Sequence P - Raw Data

Sequence P: Wake Flow Visualization, Upwind (P) This test sequence used an upwind, rigid turbine with a 0° cone angle. The wind speed ranged from 5 m/s to 15 m/s. Yaw angles of 0° to –60° were achieved. The blade tip pitch was 3°. The rotor rotated at 72 RPM. Blade and probe pressure measurements were collected. The teeter dampers were replaced with rigid links, and these two channels were flagged as not applicable by setting the measured values in the data file to –99999.99 Nm. The teeter link load cell was pre-tensioned to 40,000 N. The aluminum blade tip designed to contain a smoke generator was installed, and counterweights were installed in the non-instrumented blade tip to compensate. The turbine was positioned at the appropriate yaw angle, and the smoke generator was ignited remotely. The campaign duration was 3 minutes for all tests except P1000000, which was 2 minutes. The file name convention was the standard format except for P10000A0, which indicated a 3° pitch angle. File P1000000 used a 12° pitch angle. After these two campaigns were collected, it was determined that all subsequent data should be collected with a 3° pitch angle. Pressure data were not acquired during this sequence, so all associated data values were flagged as not applicable by setting the measured values in the data file to 0.000 Pa. Corresponding pressure data are available from Sequence H for the 3° pitch angle test points. Flow visualization data obtained from wall- and ceiling-mounted video cameras were recorded to videotape. The camera locations and calibration procedures are described in Appendix J.

17 WIND ENERGY↗

Transforming Drainage Research Data (USDA-NIFA Award No. 2015-68007-23193)

This dataset contains research data compiled by the “Managing Water for Increased Resiliency of Drained Agricultural Landscapes” project a.k.a. Transforming Drainage. This project was funded from 2015-2021 by the United States Department of Agriculture, National Institute of Food and Agriculture (USDA-NIFA, Award No. 2015-68007-23193). Data are also available from a separate web-accessible application (drainagedata.org). At drainagedata.org, users can visualize the data with customized tools, query based on specific sites and measurements of interest, and access site photographs, maps, summaries, and publications. Additional data or edits made following the publication of this data here at USDA NAL Ag Data Commons will be posted under the Versions tab on drainagedata.org. These data began in 1996 and include plot- and field-level measurements for 39 experiments across the Midwest and North Carolina. Practices studied include controlled drainage, drainage water recycling, and saturated buffers. In total, 219 variables are reported and span 207 site-years for tile drainage, 154 for nitrate-N load, 181 for water quality, 92 for water table, and 201 for crop yield.

Modeling↗

Sensitivity Analysis of Drivers Water Shortage in the Los Angeles Region During Drought

The code and detailed step-by-step instructions for generating the model output data, processing results, and analysis and plotting are provided at https://github.com/IMMM-SFA/Ferencz_et_al_2026_ER_Water. The PyArtes model is a python adaptation of the Artes model. PyArtes uses many of the same input data and optimization model architecture as Artes. Documentation for the PyArtes model is provided in the Supplement to the paper. The primary data product are simulated monthly water shortages for indoor and outdoor demand under a large ensemble of drought scenarios (>13,000). The droughts are hypothetical and are not based on historical time series data of supply sources - though historical data did help inform ranges explored for supply parameters. Demands are informed by recent 2017-2021 water supply data. Demands used for the model can be accessed at https://github.com/IMMM-SFA/Ferencz_et_al_2026_ER_Water. Simulations resolve demand for over 90 water providers in the study region. The results report 36 months of water shortage data for each indoor and outdoor demand node. The study also developed a multilayer perceptron (MLP) neural network trained on a subset of the simulated shortage ensemble to emulate worst annual water shortage for a given set of parameter multipliers -- provided the parameter values fall within the ranges sampled in the ensemble. Emulated water shortages for synthetic ensembles are in the MLP-generated shortages folder. The MLP model was used to generate larger ensembles to support Sobol analysis that would have been extremely computationally expensive to simulate. Datasets provided in this repository*: Simulated shortages. These results are used for the analysis for Figures 5, 8, and 9 in the paper, and also to train the MLP emulator. .zip file containing outputs for the 13,312 scenario ensemble. Separate .csv files for indoor and outdoor shortage for each scenario. Rows = demand ids (~100), Columns = months (36) Units = acre-feet/month of shortage (shortage = monthly demand - supply). 1 acft = 1233.48 m^3 .csv files of aggregated shortages derived from the 13,312 ensemble Rows = scenarios (13,312), Columns = demand ids (~100) Units = acre-feet/year (either worst annual shortage or total shortage over the 3-year drought) .csv file of the parameter multipliers scenarios for the ensemble .csv file of the parameter ranges and baseline values the multipliers were applied to MLP-generated shortages. These results are used for Figures 4, 6, and 7 in the paper. mwd higher folder: scenario ensembles, emulated worst year total shortages (acft), and Sobol results Emulated shortages. Rows = scenarios, columns = demand ids, units acft Sobol results. Rows = demand ids, columns Sobol (S1, ST, or 95% confidence interval) value for each parameter mwd lower folder: scenario ensembles, emulated worst year total shortages (acft), and Sobol results same organization as mwd higher MLP performance: performance metrics (R^2, RMSE, BIAS, MAPE) for the testing subset (20% or 2,662 scenarios) and simulated vs emulated worst year shortage (acre-feet/year) for every demand node, MWD wholesale regions, and the entire study region (LAC). Supporting data for figures. Figure plotting scripts in the associated GitHub repo. These files support analysis and visualization. Geospatial Data used for plotting simulated water shortages and Sobol results. Dictionary of full names for demand nodes in the model and estimates of water supply by source type informed by Artes input files and California Urban Water Management Planning data: https://water.ca.gov/Programs/Water-Use-And-Efficiency/Urban-Water-Use-Efficiency/Urban-Water-Management-Plans *Readme files provided for each folder.

drought↗

MAGIC: M arching Cubes Isosurface Uncertainty Visualization for G auss i an Uncertain Data With Spatial C orrelation

Here, in this paper, we study the propagation of data uncertainty through the marching cubes algorithm for isosurface visualization for correlated uncertain data. Consideration of correlation has been shown paramount for avoiding errors in uncertainty quantification and visualization in multiple prior studies. Although the problem of isosurface uncertainty with spatial data correlation has been previously addressed, there are two major limitations to prior treatments. First, there are no analytical formulations for uncertainty quantification of isosurfaces when the data uncertainty is characterized by a Gaussian distribution with spatial correlation. Second, as a consequence of the lack of analytical formulations,existing techniques resort to a Monte Carlo sampling approach, which is expensive and difficult to integrate into visualization tools. To address these limitations, we present a closed-form framework to efficiently derive uncertainty in marching cubes level-sets for Gaussian uncertain data with spatial correlation (MAGIC). To derive closed-form solutions, we leverage the Hinkley's derivation on the ratio of Gaussian distributions. With our analytical framework, we achieve a significant speed-up and enhanced accuracy of uncertainty quantification over classical Monte Carlo methods. We further accelerate our analytical solutions using many-core processors to achieve speed-ups up to 585× and integrability with production visualization tools for broader impact. We demonstrate the effectiveness of our correlation-aware uncertainty framework through experiments on meteorology, urban flow, and astrophysics simulation datasets.

Gaussian↗

Visualization Quality Assessment

Understanding how inaccuracies in visualizations affect users’ perception and understanding of scientific data is hard. Inaccuracies in visualizations are quite common and could arise from a range of sources such as errors in the original dataset arising from compression artifacts, errors in the capturing device, noise during transmission of the data, effects due to the algorithm being used to convert data to visualization images, images generated from neural networks, and sources we have yet to discover. Many image quality assessment metrics have been developed to quantify image errors. However, these are usually focused on “natural images” rather than visualizations of scientific data. Common image quality assessment metrics (IQAs) include MSE, PSNR, perceptual metrics such SSIM, FSIM as well as perceptual metrics using deep learning approaches. However, a critical part of understanding how errors are perceived by humans, and subsequently developing more accurate quality assessment metrics, is through user evaluation studies. The goal of this software is to develop a visualization quality assessment (VQA) process that will enable the generation of VQAs that can be used to quantify errors in scientific data visualizations. The VQA development process will include software to support user evaluation experimental design, analysis of visualization differences against standard quality metrics, and the ability to develop additional VQA metrics specific to scientific visualization images.

Grosset, Andre↗

1000 Soils Pilot Dataset, version 8, May 2025

This record hosts data generated by the 1000 Soils Pilot. Data will be updated as more become available. Please see the most recent data upload for current data. A beta visualization tool is available for some data types at https://shinyproxy.emsl.pnnl.gov/app/1000soils. Please submit any suggestions or comments through the 'contact' tab. We are actively working to improve visualizations and value all feedback. Data completed include: Geochemistry, texture, respiration, and enzyme activities FTICR-MS organic matter chemistry Microbial biomass C and N TOC/TDN of water-extractable OM X-ray computed tomography (derived metrics available here, raw data available upon request) Metagenomes; a variety of data formats are available upon request Soil hydraulic properties Data in progress: LC-MS/MS in development, timeline TBD, inquire for status 1000S_processed_BGC_summary.csv contains all available biogeochemical data; microbial biomass C and N; and TOC/TDN of water-extractable OM; and 1000S_Tomography.xslx contains a summary of data generated via X-ray computed tomography. icr_v2_corems2.csv contains FTICR-MS data processed by CoreMS version 2. These data are merged by formula across instrument runs to enable cross-sample comparisons. Technical replicates are merged by retaining peaks present in 2 out of 3 replicates. 1000Soils_Metadata_Site_Mastersheet_v1.csv contains site information. Soil Hydraulics_corrected_02042025.xlsx contains soil hydraulics information. Readme File_v4.xlsx is the readme file. Please contact the MONet project (monet.emsl@pnnl.gov) or Emily Graham (emily.graham@pnnl.gov) with questions. The following file and all raw data are available upon request: icr_by_mass_for_single_sample_analysis_only.csv contains FTICR-MS data processed by CoreMS and is intended for usage in the calculation of biochemical transformations within samples only. These data are not acceptable for cross-sample comparison of masses because they are from multiple instrument runs. For more information, please see: https://www.emsl.pnnl.gov/monet and https://sc-data.emsl.pnnl.gov/monet Acknowledgment: Soil data were provided by the Molecular Observation Network (MONet) at the Environmental Molecular Sciences Laboratory (https://ror.org/04rc0xn13), a DOE Office of Science user facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830. The work (proposal: 10.46936/10.25585/60008970) conducted by the U.S. Department of Energy, Joint Genome Institute (https://ror.org/04xm1d337), a DOE Office of Science user facility, is supported by the Office of Science of the U.S. Department of Energy operated under Contract No. DE-AC02-05CH11231. The Molecular Observation Network (MONet) database is an open, FAIR, and publicly available compilation of the molecular and microstructural properties of soil. Data in the MONet open science database can be found at https://sc-data.emsl.pnnl.gov/.

biogeochemistry↗

Continuous Emulation and Multiscale Visualization of Traffic Flow Using Stationary Roadside Sensor Data

With the advent of the next-generation traffic monitoring systems, there has been a significant increase in the spatial-temporal resolution of vehicle mobility data in many cities. Effective analysis and visualization of such data can provide transportation planners with data-driven insights, which can facilitate the understanding of multiscale traffic dynamics. In this paper, we present a web-based traffic emulator for emulating and visualizing near-real-time and historical traffic flows on highways using data from road-side sensors. To construct a continuous traffic flow, the emulator adopts an analytical pipeline that can (a) integrate traffic data collected from discrete road-side radar detection sensors, (b) interpolate traffic conditions (vehicle speed and volume) on unmeasured road segments based on traffic flow theory, and (c) generate lane-specific vehicle trajectories and movements using a mathematically optimized representation of the road network. Our app also provides an integrated visual workflow that allows users to explore the interconnected traffic dynamics using an appropriate traffic flow visualization selected based on the level of detail. We devise two innovative geo-visualization techniques that utilize an animated strips-network representation and a lane usage matrix to visualize lane performances. To ensure a smooth emulation of large-scale traffic flow in an easy-to-access web environment, we implement the emulator using client-side GPU-accelerated techniques. Lastly, we close with a case study that visualizes traffic dynamics of two scenarios - an afternoon peak hour and a traffic accident - in Chattanooga, Tennessee. Our app visualizes the responses of traffic dynamics during different traffic conditions, and to the presence of the traffic accident at different spatial scales.

42 ENGINEERING↗