Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “analysis and visualization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A Processing and Analytics System for Microscopy Data Workflows: The Pycroscopy Ecosystem of Packages

Major advancements in fields as diverse as biology and quantum computing have relied on a multitude of microscopy techniques. Despite the considerable proliferation of these instruments, significant bottlenecks remain in terms of processing, analysis, storage, and retrieval of the acquired datasets. Aside from lack of file standards, individual domain-specific analysis packages are often disjoint from the underlying datasets, and thus keeping track of analysis and processing steps remains tedious for the end-user, hampering reproducibility. Here, in this study, the pycroscopy ecosystem of packages is introduced, an open-source python-based ecosystem underpinned by a common data model. The data model, termed the N-dimensional spectral imaging data format, is realized in pycroscopy's sidpy package. This package is built on top of dask arrays, thus leveraging dask array attributes, but expanding them to accelerate microscopy relevant analysis and visualization. Several examples of the use of the pycroscopy ecosystem to create workflows for data ingestion and analysis of scanning transmission electron microscopy (STEM) and scanning probe microscopy data are shown. Adoption of such standardized routines will be critical to usher in the next generation of autonomous instruments where processing, computation, and meta-data storage will be critical to overall experimental operations.

97 MATHEMATICS AND COMPUTING↗

py4DSTEM: A Software Package for Four-Dimensional Scanning Transmission Electron Microscopy Data Analysis

Scanning transmission electron microscopy (STEM) allows for imaging, diffraction, and spectroscopy of materials on length scales ranging from microns to atoms. By using a high-speed, direct electron detector, it is now possible to record a full two-dimensional (2D) image of the diffracted electron beam at each probe position, typically a 2D grid of probe positions. These 4D-STEM datasets are rich in information, including signatures of the local structure, orientation, deformation, electromagnetic fields, and other sample-dependent properties. However, extracting this information requires complex analysis pipelines that include data wrangling, calibration, analysis, and visualization, all while maintaining robustness against imaging distortions and artifacts. In this paper, we present py4DSTEM, an analysis toolkit for measuring material properties from 4D-STEM datasets, written in the Python language and released with an open-source license. We describe the algorithmic steps for dataset calibration and various 4D-STEM property measurements in detail and present results from several experimental datasets. We also implement a simple and universal file format appropriate for electron microscopy data in py4DSTEM, which uses the open-source HDF5 standard. We hope this tool will benefit the research community and help improve the standards for data and computational methods in electron microscopy, and we invite the community to contribute to this ongoing project.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

VTK-m v. 2.0

VTK-m is a toolkit of scientific visualization algorithms for emerging processor architectures. VTK-m supports the fine-grained concurrency for data analysis and visualization algorithms required to drive extreme scale computing by providing abstract models for data and execution that can be applied to a variety of algorithms across many different processor architectures.

Moreland, Kenneth [Oak Ridge National Lab. (ORNL),↗

Online Analytics for Remedy Support at DOE Environmental Management Sites

Environmental data is important for managing environmental restoration/waste site remediation, planning of monitoring efforts, addressing climate resilience, and engaging with stakeholders and regulators. A major challenge is how to manage the many different types and the large volume of environmental data in a way that allows practitioners and site managers to understand data implications and support decisions. The Suite Of Comprehensive Rapid Analysis Tools for Environmental Sites (SOCRATES, https://www.pnnl.gov/projects/socrates) is a web application that provides data access, visualization, and rapid analytics to help make sense of environmental data, support remedy decisions, and communicate information. Development of SOCRATES has been funded through the DOE Richland Operations Office (RL) to support communication and decision making for the Hanford Site, thus is only tied into Hanford environmental data. However, the capabilities of SOCRATES are more broadly applicable to DOE-EM sites engaged in environmental remediation and management. This report describes the work to develop mechanisms for bringing non-Hanford data into SOCRATES so that other DOE-EM sites could make use of the visualization and analysis capabilities to support communication and decision making related to managing environmental restoration/waste site remediation, optimization/exit strategies for pump-and-treat systems, planning monitoring efforts, addressing climate resilience, and/or engaging with stakeholders and regulators. The background, approach, data transfer formats, examples, and next steps for this new SOCRATES-EM software are described in this report.

54 ENVIRONMENTAL SCIENCES↗

Guiding the choice of informatics software and tools for lipidomics research applications

Progress in mass spectrometry lipidomics has led to a rapid proliferation of studies across biology and biomedicine. These generate extremely large raw datasets requiring sophisticated solutions to support automated data processing. To address this, numerous software tools have been developed and tailored for specific tasks. However, for researchers, deciding which approach best suits their application relies on ad hoc testing, which is inefficient and time consuming. Here we first review the data processing pipeline, summarizing the scope of available tools. Next, to support researchers, LIPID MAPS provides an interactive online portal listing open-access tools with a graphical user interface. This guides users towards appropriate solutions within major areas in data processing, including (1) lipid-oriented databases, (2) mass spectrometry data repositories, (3) analysis of targeted lipidomics datasets, (4) lipid identification and (5) quantification from untargeted lipidomics datasets, (6) statistical analysis and visualization, and (7) data integration solutions. Detailed descriptions of functions and requirements are provided to guide customized data analysis workflows.

59 BASIC BIOLOGICAL SCIENCES↗

VTK-m: Visualization for the Exascale Era and Beyond

A recent trend in modern high-performance computing is the increasing use of hybrid architectures, where the vast majority of performance comes from accelerators. Modern accelerators are based on Graphics Processing Units (GPU) that contain many low power cores that in their aggregate provides an extremely high computation rate. Current and future CPU processors are requiring more explicit parallelism as each successive version of the hardware packs in more cores, and technologies like hyperthreading and vector operations require even more parallel processing to leverage each core’s full potential. As an example, the Frontier supercomputer installed at Oak Ridge National Laboratories recently hit a record breaking 1.1 exaflops1 on the LINPACK HPC benchmark [Shoemaker 2022]. The system contains 37632 AMD MI250x GPUs which requires more than half a billion threads to keep the system fully utilized [Khizeran 2022].VTK-m is a toolkit of scientific visualization algorithms for these emerging processor architectures. VTK-m supports the fine-grained concurrency for data analysis and visualization algorithms required to drive extreme scale computing by providing abstract models for data and execution that can be applied to a variety of algorithms across many different processor architectures.

Bolstad, Mark↗

COVID-19 Data Curation Effort: An Initial Analysis of the Data

During the COVID-19 pandemic of 2020, major case reporting outlets quickly coalesced around two or three primary vendors. Johns Hopkins University and The New York Times were among the more prominent, and all were of great value to the nation, particularly during the uncertain early stages of the pandemic. They primarily focused on three major attributes: number of new cases, deaths, and recovery. Recognizing that many states were reporting very detailed data sets (e.g., hospital beds) at a county level or finer, the ORNL Pandemic Modeling team embarked on a major data curation effort from March to June 2020 for the purpose of capturing this wealth of detailed data. The challenge of curating this data was daunting. The number of attributes reported by the states grew on almost on a weekly basis. States were routinely shifting their web tool strategies away from easily parsable HTML-based formatting to new Tableau and ArcGIS content. This growth in the sheer number of attributes, combined with the unpredictable shifts in data format, meant an aggressive and agile combination of automated scripting and manual scraping was required to capture new daily streams. Further, the team had to scale up staff and widen its approach for capture and storage. As a result, the team collected more than 11 million data points. Following the close of this data collection effort on June 30 th , 2020, the team embarked on a major effort to appraise what had been collected, including an inventory list, spatial completeness, temporal completeness, scale and geographic characteristics, and a determination. A report on this matter was submitted on September 15 th , 2020, titled “DOE COVID-19 Data Curation Effort: Overview of Data Collection Coverage”. Over 2000 unique attributes had been netted over a wide range of spatial scales, including state, county, zip codes, health regions, and census blocks. Over 11 million individual data points were collected across these attributes, and spatial coverage (in total) included all 50 states and multiple territories. What became apparent in the process is that in the absence of any data standards, many states reported a wide variety of unique attributes that were not always compatible with attributes reported in other states. As time continued, states began adding new attributes and offering finer grain detail in some older attributes. This meant that not all data streams existed for the entire time period; in fact, the number tended to increase dramatically towards the end. Often, states would begin an attribute series and then stop altogether. These highly variable and uncertain conditions illuminated the need for harmonization approaches that would reconcile and conflate changing attribute names and detail over time. For example, grouping racial data reported as either Black or African American, depending on the state, into a single harmonized attribute. These choices would make a within-state analysis possible during the time period and lead to potential between-state analytics later on. This was almost entirely a manual decision process, requiring some subjective decision-making at times, to prevent a fragmented, short-lived collection of time series fragments that would offer few insights into trends, patterns, and correlates. This report imports harmonized data for state and county into the World Spatio-Temporal Analytics and Mapping Project (WSTAMP). WSTAMP is a major space-time analysis and visualization tool developed at ORNL for the National Geospatial-Intelligence Agency specifically for this kind of exploratory analysis. WSTAMP offers a rich analytical and graphical environment consisting of a wide range of analytics. These include time series plots, statistical summaries, data mining techniques, trend and pattern detection, and hypothesis generation.

59 BASIC BIOLOGICAL SCIENCES↗

Top Research Challenges and Opportunities for Near Real-Time Extreme-Scale Visualization of Scientific Data

The rapid advancement in scientific simulations and experimental facilities has resulted in the generation of vast amounts of data at unprecedented scales. The analysis and visualization of large amounts of data is a challenge in and of itself, but the requirements for timeliness significantly magnify these difficulties. Near real-time visualization is critical to monitor and analyze the data produced by these large facilities, but current production tools are not well-suited to these requirements. In this position paper, we share our perspective on some of the challenges, and thus, opportunities for research that stand in the way of near-real-time visualization of large scientific data.

Pugmire, Dave↗

Remote Instrumentation and Data Acquisition

This poster outlines the development and implementation of a remote data acquisition system for waveform analysis using a Rohde & Schwarz oscilloscope. The project involved capturing waveform data, and transferring it to a local machine for visualization and analysis. The core logic was developed in C++ with a focus on object oriented programming and the use of polymorphism so the main application can interact with any instrument without knowing its exact type, simplifying the overall logic and making it easier to add or swap out components without changing the rest of the codebase.. The system issues Standard Commands for Programmable Instruments (SCPI) via a socket connection and parses the oscilloscope s ASCII waveform data. The C++ application was containerized using Docker for ease of portability, and reproducibility. Emphasis was placed on secure networking practices, error handling, and effective data capture. The report describes the technical steps taken, challenges encountered, and future work, providing insight into the practical integration of hardware interfacing with remote computational environments.

Parikh, Jaymil [Illinois U., Urbana]↗

Remote Instrumentation and Data Acquisition: An Internship Research Report

This report outlines the development and implementation of a remote data acquisition system for waveform analysis using a Rohde & Schwarz oscilloscope. The project involved capturing waveform data, and transferring it to a local machine for visualization and analysis. The core logic was developed in C++ with a focus on object oriented programming and the use of polymorphism so the main application can interact with any instrument without knowing its exact type, simplifying the overall logic and making it easier to add or swap out components without changing the rest of the codebase.. The system issues Standard Commands for Programmable Instruments (SCPI) via a socket connection and parses the oscilloscope’s ASCII waveform data. The C++ application was containerized using Docker for ease of portability, and reproducibility. Emphasis was placed on secure networking practices, error handling, and effective data capture. The report describes the technical steps taken, challenges encountered, and lessons learned, providing insight into the practical integration of hardware interfacing with remote computational environments.

Parikh, Jaymil [Fermilab]↗

Slycat Enables Synchronized 3D Comparison of Surface Mesh Ensembles [Brief]

In support of analyst requests for Mobile Guardian Transport studies, researchers at Sandia National Laboratories have expanded data types for the Slycat ensemble-analysis and visualization tool to include 3D surface meshes. This new capability represents a significant advance in our ability to perform detailed comparative analysis of simulation results. Analyzing mesh data rather than images provides greater flexibility for post-processing exploratory analysis.

36 MATERIALS SCIENCE↗

Root system architecture and environmental flux analysis in mature crops using 3D root mesocosms

Current methods of root sampling typically only obtain small or incomplete sections of root systems and do not capture their true complexity. To facilitate the visualization and analysis of full-sized plant root systems in 3-dimensions, we developed customized mesocosm growth containers. While highly scalable, the design presented here uses an internal volume of 45 ft 3 (1.27 m 3 ), suitable for large crop and bioenergy grass root systems to grow largely unconstrained. Furthermore, they allow for the excavation and preservation of 3-dimensional root system architecture (RSA), and facilitate the collection of time-resolved subterranean environmental data. Sensor arrays monitoring matric potential, temperature and CO 2 levels are buried in a grid formation at various depths to assess environmental fluxes at regular intervals. Methods of 3D data visualization of fluxes were developed to allow for comparison with root system architectural traits. Following harvest, the recovered root system can be digitally reconstructed in 3D through photogrammetry, which is an inexpensive method requiring only an appropriate studio space and a digital camera. We developed a pipeline to extract features from the 3D point clouds, or from derived skeletons that include point cloud voxel number as a proxy for biomass, total root system length, volume, depth, convex hull volume and solidity as a function of depth. Ground-truthing these features with biomass measurements from manually dissected root systems showed a high correlation. We evaluated switchgrass, maize, and sorghum root systems to highlight the capability for species wide comparisons. We focused on two switchgrass ecotypes, upland (VS16) and lowland (WBC3), in identical environments to demonstrate widely different root system architectures that may be indicative of core differences in their rhizoeconomic foraging strategies. Finally, we imposed a strong physiological water stress and manipulated the growth medium to demonstrate whole root system plasticity in response to environmental stimuli. Hence, these new “3D Root Mesocosms” and accompanying computational analysis provides a new paradigm for study of mature crop systems and the environmental fluxes that shape them.

59 BASIC BIOLOGICAL SCIENCES↗

From 2D to 4D: a containerized workflow and browser to explore dynamic chromatin architecture

Background Characterizing the physical organization of the genome is essential for understanding long-range gene regulation, chromatin compartmentalization, and epigenetic accessibility. Hi-C experiments generate two-dimensional (2D) genome-wide contact maps of chromatin interactions by capturing the spatial proximity between genomic loci, which reveal interaction frequencies but lack the spatial resolution needed to interpret the three-dimensional (3D) genome structure(s). Emerging evidence suggests that epigenetic regulation is closely linked to 3D genome architecture, and that structural changes over time (4D) drive key biological processes in development, disease, and environmental response. Thus, integrating 3D structure with functional data is critical for a more complete understanding of genome regulation. Previous work, most notably the 4DHiC chromosome modeling framework, has shown that physical multi-dimensional modeling approaches rooted in polymer physics and molecular dynamics can resolve these structures at biologically meaningful resolutions by integrating temporal Hi-C data with physical constraints to uncover dynamic chromosome reorganization. Thus, molecular dynamics simulations, constrained by Hi-C contact matrices, can resolve fine-scale structural changes and reveal functionally significant transitions in chromatin conformation. Results Herein, we present the 4D Genome Browser Workflow (4DGBWorkflow) and the 4D Genome Browser (4DGB). The algorithm is based on the 4DHiC method, and the containerized tool is an end-to-end workflow that can transform, filter, and view 4D epigenomics and chromatin datasets, allowing non-specialists to apply three-dimensional modeling principles to diverse datasets and experimental conditions. The software executes on a laptop running macOS, Linux or Windows. From input Hi-C files (.hic), the 4DGBWorkflow produces 3D reconstructions of chromosomes, integrates the reconstruction with track data (e.g., epigenetic marks, transcriptome profiles), and provides comparative visualization of the results in a single workflow. Conclusions The 4DGBWorkflow and 4D Genome Browser are open-source tools for comparative analysis and visualization of 4D chromosome datasets, including chromatin architecture and epigenomic signals. Automatic integration of Hi-C data with molecular dynamics democratizes the construction of time resolved 3D genome structures, simplifying complex simulations and data integration schemes.

3D Genome Browser↗

Integrated, Interoperable Software Environment for Fusion Simulation and Data Analysis Tools SBIR Phase II

This SBIR effort was focused on developing a production ready system to address the integration and interoperability challenges with analysis and visualization in fusion simulations. Our overarching technical objective was to minimize the code development simulation scientists incur when coupling their simulation codes with different analysis frameworks. To this end, we developed an open-source software library to make data exchange between application easier and a Web application to manage and display analysis extracts from simulations. We have also augmented existing libraries funded by DOE such as ADIOS and VTK-m. When used together, these make it significantly easier to integrate simulation and analysis capability. We demonstrated the flexibility of our approach using two common simulation codes in the fusion community, XGC1 and GTC.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Open Reproducible Electron Microscopy Data Analysis

Electron microscopy (EM) is a cornerstone technique in the materials and biological sciences capable of imaging structures at nano- to atomic-scale resolution. Advances in technologies mean that one acquires datasets at increasing data rates and sizes. These advancements present enormous opportunities for researchers to understand complex systems. However, processing the resulting large-scale, complex data in a reproducible and shareable way is a real challenge for researchers. The building, managing, and maintaining complex workflows in a reproducible manner requires extensive knowledge in several areas outside the researchers’ core skill sets, such as software engineering, data science, and high-performance computing (HPC). Our work demonstrates an innovative approach to solving these problems, enabling reproducible EM data analysis through container encapsulated pipelines. Using modern container technologies, we encapsulate processing elements and connect them using shared memory. We expose user-friendly, advanced algorithms and tools to allow end users to utilize without expert programming skills. The platform enables reproducible, scalable, shareable pipelines for the analysis and visualization of EM data. Focusing on interoperability, we leverage the DOE and other agencies’ existing investments to provide a powerful software platform for EM data analysis.

Harris, Christopher↗

Continuous Emulation and Multiscale Visualization of Traffic Flow Using Stationary Roadside Sensor Data

With the advent of the next-generation traffic monitoring systems, there has been a significant increase in the spatial-temporal resolution of vehicle mobility data in many cities. Effective analysis and visualization of such data can provide transportation planners with data-driven insights, which can facilitate the understanding of multiscale traffic dynamics. In this paper, we present a web-based traffic emulator for emulating and visualizing near-real-time and historical traffic flows on highways using data from road-side sensors. To construct a continuous traffic flow, the emulator adopts an analytical pipeline that can (a) integrate traffic data collected from discrete road-side radar detection sensors, (b) interpolate traffic conditions (vehicle speed and volume) on unmeasured road segments based on traffic flow theory, and (c) generate lane-specific vehicle trajectories and movements using a mathematically optimized representation of the road network. Our app also provides an integrated visual workflow that allows users to explore the interconnected traffic dynamics using an appropriate traffic flow visualization selected based on the level of detail. We devise two innovative geo-visualization techniques that utilize an animated strips-network representation and a lane usage matrix to visualize lane performances. To ensure a smooth emulation of large-scale traffic flow in an easy-to-access web environment, we implement the emulator using client-side GPU-accelerated techniques. Lastly, we close with a case study that visualizes traffic dynamics of two scenarios - an afternoon peak hour and a traffic accident - in Chattanooga, Tennessee. Our app visualizes the responses of traffic dynamics during different traffic conditions, and to the presence of the traffic accident at different spatial scales.

42 ENGINEERING↗

A quantitative comparison of the fingerprint of twinned microstructures through surface and three-dimensional techniques

Assessing the fingerprint of a material’s microstructure is key for supporting materials design. With the emergence of a wide range of 3D characterization techniques, it is critical to understand the main differences in fingerprints reconstructed from 2D and 3D datasets. To this end, we introduce a graph-based microstructure reconstruction framework that enables structural comparisons of twin domain networks in high purity Ti using 3D and 2D electron backscatter diffraction. Insights into the structure of the twin networks are facilitated by combining statistical analysis of twin crystallography with visual and graphical analysis of the novel graph abstractions of the twins. We demonstrate that compared to 3D reconstructions, conventional 2D views of twinning miss key aspects of the microstructure including the high interconnectivity of domains into networks that span the full reconstruction volume. The reduced cross-grain and in-grain twin connectivity typically observed in 2D has notable implications on our understanding of how twinning mediates the plastic response of microstructures and how twin networks evolve. It is thus clear that 3D characterization is critical for accurately inferring both twin network morphologies as well as the key unit processes facilitating network formation.

36 MATERIALS SCIENCE↗