Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scientific data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

SIRIUS: Science-Driven Data Management for Multi-Tiered Storage

The data sets being generated by large applications on very large-scale systems are increasing in both size and complexity. At the same time, there are new ways available to store and access these data sets. The goal in this project is to develop software that applications can use to make use of new and existing storage technologies in more sophisticated ways. One challenge in scientific data management is handling ‘hot’ vs ‘cold’ data. Data that is hot is data that is needed (or will be needed soon) in order for the program to continue progressing, while cold data is either output (and so will not be need further during the life of the program) or will not be needed until significantly later in the program’s run. Hot data should be stored in a way that allows fast access. On most systems, economic factors lead to an inverse relationship between storage performance and storage capacity and so fast access storage is limited. This makes it important to correctly place hot and cold data and avoid cold data unnecessarily consuming precious resources. In this reporting period, we addressed this challenge in various ways and at various levels. Data management frameworks offer only limited control to applications in how data is stored. We have added software capabilities for seamlessly moving data between layers of the storage technology using promote and demote functions to existing software frameworks. This gives direct control to applications in deciding what priority data receives. Additionally, we integrated different storage layer management frameworks in order to allow data to be exchanged and moved between storage layers in a consistent way across the application. Further, applications are not always able to directly decide what storage level makes sense for a given piece of data without an understanding of the underlying storage technologies. Data storage frameworks are often positioned to make these sorts of decisions in service of the application. We have added machine-learning based capabilities to data staging frameworks in order to make intelligent decisions about where data should be stored given learning about patterns in previous usage of similar data.

97 MATHEMATICS AND COMPUTING↗

SPRUCE Vegetation Phenology in Experimental Plots from Phenocam Imagery, 2015-2022

This data set consists of PhenoCam data from the SPRUCE experiment from the beginning of whole ecosystem warming (Hanson et al. 2017) in August 2015 through March 31 of 2023, with start- and end-of-season phenological transition dates derived through the end of autumn 2022. Digital cameras, or phenocams, installed in each SPRUCE enclosure track seasonal variation in vegetation “greenness”, a proxy for vegetation phenology and associated physiological activity. Three separate regions of interest (ROIs) were defined for each camera field of view, corresponding to different vegetation types and demarcating (1) Picea trees (vegetation type EN, for evergreen needleleaf); (2) Larix trees (vegetation type DN, for deciduous needleleaf); and (3) the mixed shrub layer (vegetation type SH). This data set consists of three sets of data files: 3-day summary product files: One file for each camera and each ROI (i.e. vegetation type), characterizing vegetation color at a 3-day time step Contains 35 files in *.csv format inside a compressed (*.zip) file. Transition date file: Estimates “greenness rising” (spring) and “greenness falling” (autumn) transition dates derived from the smoothed daily green chromatic coordinate (GCC) values, for each camera and each ROI (i.e. vegetation type) Contains one file in *.csv format Snow flag files: Indicate days with snow on trees or snow on ground for each experimental enclosure Contains two files in *.csv format, one for snow on trees and one for snow on ground This data set consists of two sets of companion files: Accompanying HTML files show the 90th quantiles of the mean GCC plotted together with transition dates for each vegetation type and plot. Contains three files in HTML format, one for each vegetation type One additional file in HTML format with the transition dates plotted for each vegetation type, by year R files for processing Phenocam files and flags. Contains five files in R file (*.R) format in one compressed (*.zip) file User Note: All imagery is posted in near-real time to the PhenoCam Project web page (http://phenocam.sr.unh.edu/), where it is publicly available. Scroll to “spruce” in the Gallery or link directly to the 29 SPRUCE cameras at https://tinyurl.com/y7z5mau7. The data reported here are based on the complete camera record from SPRUCE and supersedes the previously released phenocam datasets (see Related Data Sets). The estimated transition dates for previously released datasets may differ slightly (in most cases, by ±3 days or less), because following standard PhenoCam processing protocols (Richardson et al. 2018, Scientific Data), smoothing and interpolation, outlier removal, and transition date estimation are always conducted using the full data record.

SPRUCE Experiment, Marcell Experimental Forest, Sp↗

SPRUCE Vegetation Phenology in Experimental Plots from Phenocam Imagery, 2015-2023

This data set consists of PhenoCam data from the SPRUCE experiment from the beginning of whole ecosystem warming (Hanson et al. 2017) in August 2015 through March 31 of 2024, with start- and end-of-season phenological transition dates derived through the end of autumn 2023. Digital cameras, or phenocams, installed in each SPRUCE enclosure track seasonal variation in vegetation “greenness”, a proxy for vegetation phenology and associated physiological activity. Three separate regions of interest (ROIs) were defined for each camera field of view, corresponding to different vegetation types and demarcating (1) Picea trees (vegetation type EN, for evergreen needleleaf); (2) Larix trees (vegetation type DN, for deciduous needleleaf); and (3) the mixed shrub layer (vegetation type SH). This data set consists of three sets of data files: (1) 3-day summary product files: One file for each camera and each ROI (i.e. vegetation type), characterizing vegetation color at a 3-day time step • Contains 36 files in *.csv format inside a compressed (*.zip) file. (2) Transition date file: Estimates “greenness rising” (spring) and “greenness falling” (autumn) transition dates derived from the smoothed daily green chromatic coordinate (GCC) values, for each camera and each ROI (i.e. vegetation type) • Contains one file in *.csv format (3) Snow flag files: Indicate days with snow on trees or snow on ground for each experimental enclosure • Contains two files in *.csv format, one for snow on trees and one for snow on ground This data set consists of two sets of companion files: (1) Accompanying HTML files show the 90th quantiles of the mean GCC plotted together with transition dates for each vegetation type and plot. • Contains three files in HTML format, one for each vegetation type • One additional file in HTML format with the transition dates plotted for each vegetation type, by year (2) R files for processing Phenocam files and flags. • Contains five files in R file (*.R) format in one compressed (*.zip) file User Note: All imagery is posted in near-real time to the PhenoCam Project web page (http://phenocam.sr.unh.edu/), where it is publicly available. Scroll to “spruce” in the Gallery or link directly to the 29 SPRUCE cameras at https://tinyurl.com/sprucecams. The data reported here are based on the complete camera record from SPRUCE and supersedes the previously released data inclusive of the 2015-2022 data (see Related Data Sets). The estimated transition dates for previously released datasets may differ slightly (in most cases, by ±3 days or less), because following standard PhenoCam processing protocols (Richardson et al. 2018, Scientific Data), smoothing and interpolation, outlier removal, and transition date estimation are always conducted using the full data record.

Spruce and Peatland Responses Under Changing Envir↗

SPRUCE Vegetation Phenology in Experimental Plots from PhenoCam Imagery, 2015-2024

This data set consists of PhenoCam data from the SPRUCE experiment from the beginning of whole ecosystem warming (Hanson et al. 2017) in August 2015 through March 31 of 2025 (2015-08-24 to 2025-03-31), with start- and end-of-season phenological transition dates derived through the end of autumn 2024. Digital cameras, or phenocams, installed in each SPRUCE enclosure track seasonal variation in vegetation “greenness”, a proxy for vegetation phenology and associated physiological activity. Three separate regions of interest (ROIs) were defined for each camera field of view, corresponding to different vegetation types and demarcating (1) Picea trees (vegetation type EN, for evergreen needleleaf); (2) Larix trees (vegetation type DN, for deciduous needleleaf); and (3) the mixed shrub layer (vegetation type SH). This data set consists of three sets of data files: (1) 3-day summary product files: One file for each camera and each ROI (i.e. vegetation type), characterizing vegetation color at a 3-day time step. • Contains 36 files in *.csv format inside a compressed (*.zip) file. (2) Transition date file: Estimates “greenness rising” (spring) and “greenness falling” (autumn) transition dates derived from the smoothed daily green chromatic coordinate (GCC) values, for each camera and each ROI (i.e., vegetation type). • Contains one file in *.csv format. (3) Snow flag files: Indicate days with snow on trees or snow on ground for each experimental enclosure. • Contains two files in *.csv format, one for snow on trees and one for snow on ground. This data set consists of two sets of companion files: (1) Accompanying HTML files show the 90th quantiles of the mean GCC plotted together with transition dates for each vegetation type and plot. • Contains three files in HTML format, one for each vegetation type. • One additional file in HTML format with the transition dates plotted for each vegetation type, by year. (2) R files for processing PhenoCam files and flags. • Contains five files in R file(*.R) format and the components of the phenocamr package (Version 1.1.4) used for calculating transition dates for 2015-2024. These are contained in a compressed (*.zip) file. User Note: All imagery is posted in near-real time to the PhenoCam Project web page (https://phenocam.nau.edu), where it is publicly available. Scroll to “spruce” in the Gallery or link directly to the 29 SPRUCE cameras at https://tinyurl.com/sprucecams. This data set is based on the complete camera record from SPRUCE and supersedes all previously released PhenoCam datasets (see Related Data Sets). The estimated transition dates for previously released datasets may differ slightly (in most cases, by ±3 days or less), because following standard PhenoCam processing protocols (Richardson et al. 2018, Scientific Data), smoothing and interpolation, outlier removal, and transition date estimation are always conducted using the full data record.

54 ENVIRONMENTAL SCIENCES↗

Deep Learning and Structural Imaging of Materials

Deep learning has had a transformative effect on numerous domains and is actively utilized by many scientists in data-intensive fields such as high-energy physics and cosmology. Materials science, and in particular, the structural imaging of materials with electrons and X-rays are projected to enter the age of scientific data torrents, positioning them as new application spaces for modern artificial intelligence. In this contribution, we provide a synopsis on the foundations and latest progress in deep learning and present an in-depth application of utilizing modern deep artificial neural networks in scanning transmission electron microscopy to extract structural material properties. We use this case study to expose the strengths of deep learning-based models and discuss their current limitations, in the process highlighting their potential use in other data-intensive structural imaging modalities.

Laanait, Nouamane↗

CEAZ: Accelerating Parallel I/O Via Hardware-Algorithm Co-Designed Adaptive Lossy Compression

As supercomputers continue to grow to exa-scale, the amount of data that needs to be saved or transmitted is exploding. To this end, many previous works have studied using error-bounded lossy compressors to reduce the data size and improve the I/O performance. However, little work has been done for effectively offloading lossy compression onto FPGA-based SmartNICs to reduce the compression overhead. In this paper, we propose a hardware-algorithm co-design of efficient and adaptive lossy compressor for scientific data on FPGAs (called CEAZ) to accelerate parallel I/O. Our contribution is fourfold: (1) We propose an efficient Huffman coding approach that can adaptively update Huffman codewords online based on codewords generated offline (from a variety of representative scientific datasets). (2) We derive a theoretical analysis to support a precise control of compression ratio under an error-bounded compression mode, enabling accurate offline Huffman codewords generation. This also help us create a fixed-ratio compression mode for consistent throughput. (3) We develop an efficient compression pipeline by adopting cuSZ’s dual-quantization algorithm to our hardware use case. (4) We evaluate CEAC on five real-world datasets with both a single FPGA board and 256 nodes from Bridges2 supercomputer. Experiments show that CEAZ outperforms the second-best FPGA-based lossy compressor by 2× of throughput and 9.6× of compression ratio. It also improves MPI_File_write and MPI_Gather throughputs by up to 32.7× and 31.4×, respectively.

Zhang, Chengming↗

Data Integration and Visualization for Enhanced Resilience and Sustainability in Hydropower (DIVERS-H)

U.S. hydropower plants face potential threats from shrinking water supply, rising demands, and warmer stream temperatures from various causes. Power plant owners, operators, and regulators require new tools to take advantage of and interpret the diverse range of scientific data being produced by both observational methods (for example, satellite, radar, stream gauges) and computer modeling methods that evaluate and predict how earth's dynamic systems (atmosphere, oceans, land surface, and sea ice) are changing and interacting. Combining datasets such as these with AI-based analyses introduces a novel decision support system to help users anticipate and address potential impacts on power generation stations. This new technology has been named DIVERS-H for "Data Integration and Visualization for Enhanced Resilience and Sustainability in Hydropower." In Phase I, technical feasibility was established with the development and demonstration of all the new technologies that are required. Most notably, DIVERS-H will use new artificial intelligence (AI) methods to capture the complex dynamics of water availability, demand, and environmental changes. In addition, new data management software was developed, and a prototype user interface was implemented as the precursor to a full scale decision support system. With technical research complete, the project focus now shifts to development of a commercial software product to provide users with actionable insight into water availability and the risk/resilience of critical systems at their locations of interest. Although DIVER-H was originally conceived as a tool for hydroelectric power applications, the same underlying technology can be readily applied to other water-consuming systems including coal, natural gas, oil, and nuclear power plants.

Chaudhary, Aashish [Kitware, Inc., Clifton Park, N↗

Probabilistic data fusion and physics-informed machine learning: A new paradigm for modeling under uncertainty, and its application to accelerating the discovery of new materials

In this report we summarize the work conducted by PI Perdikaris and his group under this Early Career project DE–SC0019116 during the period of 09/01/2018 – 08/31/2023. The central aim of the work was to introduce a new paradigm for scientific data analysis that can seamlessly synthesize rigorous mathematical modeling with data of variable fidelity (e.g., measurements at multiple scales/resolutions or predictions of variable fidelity models) and multiple modalities (e.g., images, time–series, or scattered measurements). The setting we are interested in involves complex systems that are partially observed and whose dynamical behavior could be hard to model or totally unknown. The inherent uncertainty associated with this setting necessitates a departure from the classical deterministic realm of modeling and scientific computation, and, consequently, our main building blocks can no longer be crisp deterministic numbers and governing laws, but instead we must operate with probabilistic models.

97 MATHEMATICS AND COMPUTING↗

5G Enabled Energy Innovation: Advanced Wireless Networks for Science

Digital wireless communication has become a foundational technology for the nation. The U.S. Department of Energy’s Office of Science (DOE-SC) is the Nation’s largest supporter of basic research in the physical sciences discovering new materials, designing advanced microelectronics, and understanding the physics of radio frequency signaling. The expanding national rollout of a new fifth-generation (5G) mobile network, coupled with the torrent of scientific data generated by next-generation devices such as battery-powered Internet of Things (IoT) sensors, has created an urgent need to enhance cutting-edge wireless technology. Breakthroughs in the deployment, integration, security, and operational range of wireless networking can provide new scientific capabilities for the next decade - from autonomous mobile instruments for scientific user facilities to intelligent sensors networks distributed over thousands of kilometers to study environmental processes. To realize this promise, however, we must continue to drive innovations in computing, artificial intelligence (AI), advanced materials, high-speed networking, and microelectronics. In March 2020, the DOE-SC convened a workshop to identify the potential opportunities and explore the scientific challenges of advanced wireless technologies.

42 ENGINEERING↗

Reliable edge machine learning hardware for scientific applications

Extreme data rate scientific experiments create massive amounts of data that require efficient ML edge processing. This leads to unique validation challenges for VLSI implementations of ML algorithms: enabling bit-accurate functional simulations for performance validation in experimental software frameworks, verifying those ML models are robust under extreme quantization and pruning, and enabling ultra-fine-grained model inspection for efficient fault tolerance. We discuss approaches to developing and validating reliable algorithms at the scientific edge under such strict latency, resource, power, and area requirements in extreme experimental environments. We study metrics for developing robust algorithms, present preliminary results and mitigation strategies, and conclude with an outlook of these and future directions of research towards the longer-term goal of developing autonomous scientific experimentation methods for accelerated scientific discovery.

Baldi, Tommaso↗

Prescreening-Based Subset Selection for Improving Predictions of Earth System Models With Application to Regional Prediction of Red Tide

We present the ensemble method of prescreening-based subset selection to improve ensemble predictions of Earth system models (ESMs). In the prescreening step, the independent ensemble members are categorized based on their ability to reproduce physically-interpretable features of interest that are regional and problem-specific. The ensemble size is then updated by selecting the subsets that improve the performance of the ensemble prediction using decision relevant metrics. We apply the method to improve the prediction of red tide along the West Florida Shelf in the Gulf of Mexico, which affects coastal water quality and has substantial environmental and socioeconomic impacts on the State of Florida. Red tide is a common name for harmful algal blooms that occur worldwide, which result from large concentrations of aquatic microorganisms, such as dinoflagellate Karenia brevis, a toxic single celled protist. We present ensemble method for improving red tide prediction using the high resolution ESMs of the Coupled Model Intercomparison Project Phase 6 (CMIP6) and reanalysis data. The study results highlight the importance of prescreening-based subset selection with decision relevant metrics in identifying non-representative models, understanding their impact on ensemble prediction, and improving the ensemble prediction. These findings are pertinent to other regional environmental management applications and climate services. Additionally, our analysis follows the FAIR Guiding Principles for scientific data management and stewardship such that data and analysis tools are findable, accessible, interoperable, and reusable. As such, the interactive Colab notebooks developed for data analysis are annotated in the paper. This allows for efficient and transparent testing of the results’ sensitivity to different modeling assumptions. Moreover, this research serves as a starting point to build upon for red tide management, using the publicly available CMIP, Coordinated Regional Downscaling Experiment (CORDEX), and reanalysis data.

54 ENVIRONMENTAL SCIENCES↗

The U.S. CMS HL-LHC R&D Strategic Plan

The HL-LHC run is anticipated to start at the end of this decade and will pose a significant challenge for the scale of the HEP software and computing infrastructure. The mission of the U.S. CMS Software & Computing Operations Program is to develop and operate the software and computing resources necessary to process CMS data expeditiously and to enable U.S. physicists to fully participate in the physics of CMS. We have developed a strategic plan to prioritize R&D efforts to reach this goal for the HL-LHC. This plan includes four grand challenges: modernizing physics software and improving algorithms, building infrastructure for exabyte-scale datasets, transforming the scientific data analysis process and transitioning from R&D to operations. We are involved in a variety of R&D projects that fall within these grand challenges. In this talk, we will introduce our four grand challenges and outline the R&D program of the U.S. CMS Software & Computing Operations Program.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Towards Interactive, Reproducible Analytics at Scale on HPC Systems

The growth in scientific data volumes has resulted in a need to scale up processing and analysis pipelines using High Performance Computing (HPC) systems. These workflows need interactive, reproducible analytics at scale. The Jupyter platform provides core capabilities for interactivity but was not designed for HPC systems. In this paper, we outline our efforts that bring together core technologies based on the Jupyter Platform to create interactive, reproducible analytics at scale on HPC systems. Our work is grounded in a real world science use case-applying geophysical simulations and inversions for imaging the subsurface. Our core platform addresses three key areas of the scientific analysis workflow-reproducibility, scalability, and interactivity. We describe our implemention of a system, using Binder, Science Capsule, and Dask software. We demonstrate the use of this software to run our use case and interactively visualize real-Time streams of HDF5 data.

containers↗

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Global Corn Heat Stress: Mean and SD of Degree Days Above 29°C based on NEX-GDDP-CMIP6 Climate Projections

Description This global dataset provides the estimated mean and standard deviation (SD) of corn heat stress (degree days above 29°C) for a set of climate models in NEX-GDDP-CMIP6 at 0.25-degree resolution. The NEX-GDDP-CMIP6 dataset is comprised of global downscaled climate scenarios derived from the General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 6 (CMIP6). The current dataset includes: Long-Term Average Degree Days Above 29°C- Historical Long-Term Average Degree Days Above 29°C- SSP245 Long-Term Standard Deviation of Degree Days Above 29°C- Historical Long-Term Standard Deviation of Degree Days Above 29°C- SSP245 The mean and SD are calculated over 1985-2014 for the historical period and over 2035-2064 for future projections. A full description of methods, including growing season, daily temperature distribution, and statistical coefficients, can be found in Haqiqi (2024). The source climate data are obtained from https://ds.nccs.nasa.gov/thredds2/catalog/catalog.html and are described in Thrasher et al (2022). The codes used to create this dataset are available at https://github.com/ihaqiqi/dd29c_nex_cmip6. Acknowledgments This work was supported by the US Department of Energy, Office of Science, Biological and Environmental Research Program, Earth and Environmental Systems Modeling, MultiSector Dynamics under Cooperative Agreement DE-SC0022141. The data processing, computation, and storage were completed on Purdue Anvil supercomputer and cyberinfrastructure supported by the National Science Foundation HDR award # 2118329: "NSF Institute for Geospatial Understanding through an Integrative Discovery Environment (I-GUIDE)". References Haqiqi. I. (2024). Trade can buffer climate-induced risks and volatilities in crop supply. Environmental Research: Food Systems. https://doi.org/10.1088/2976-601X/ad7d12 Thrasher, B., Wang, W., Michaelis, A., Melton, F., Lee, T., & Nemani, R. (2022). NASA global daily downscaled projections, CMIP6. Scientific Data, 9(1), 262. https://doi.org/10.1038/s41597-022-01393-4

Climate Change↗

Large-Scale Visualization of 3D Unstructured Groundwater Model Using Cave Automated Virtual Environment

The immersive three-dimensional (3D) virtual reality (VR) visualization of groundwater models allows us to deepen our understanding of aquifer systems and provide better solutions to present groundwater-related problems, such as groundwater recharge, water quality, and sustainability. Visualization assists in accurately developing groundwater models and revealing important subsurface features, including faulting, folding, and unconformity. However, assessing model accuracy poses challenges due to the complexity of geology and groundwater systems. This research demonstrates a workflow to visualize and analyze raw 3D unstructured groundwater model data using an immersive Cave Automated Virtual Environment (CAVE). To visualize the unstructured groundwater model data, the raw dataset is converted into interactive CAVE-compatible formats utilizing a set of tools: ParaView, Blender, and Unity. This enables researchers to immerse themselves in the data, identifying influential patterns and relationships. e resulting insights can inform the development of sophisticated machine-learning models for groundwater level prediction. The CAVE’s immersive capabilities allow intuitive exploration from various perspectives, providing a more holistic understanding of the factors affecting groundwater levels. These insights are crucial to improve predictive models. The CAVE results also facilitate collaborative analysis and have potential applications in training and education. is research demonstrates the value of immersive VR tools such as the CAVE for unraveling intricacies within high-dimensional scientific data to drive real-world forecasting and modeling applications.

54 ENVIRONMENTAL SCIENCES↗

Physically Motivated Deep Learning to Superresolve and Cross Calibrate Solar Magnetograms

Abstract Superresolution (SR) aims to increase the resolution of images by recovering detail. Compared to standard interpolation, deep learning-based approaches learn features and their relationships to leverage prior knowledge of what low-resolution patterns look like in higher resolution. Deep neural networks can also perform image cross-calibration by learning the systematic properties of the target images. While SR for natural images aims to create perceptually convincing results, SR of scientific data requires careful quantitative evaluation. In this work, we demonstrate that deep learning can increase the resolution and calibrate solar imagers belonging to different instrumental generations. We convert solar magnetic field images taken by the Michelson Doppler Imager (resolution ∼2″ pixel −1 ; space based) and the Global Oscillation Network Group (resolution ∼2.″5 pixel −1 ; ground based) to the characteristics of the Helioseismic and Magnetic Imager (resolution ∼0.″5 pixel −1 ; space based). We also establish a set of performance measurements to benchmark deep-learning-based SR and calibration for scientific applications.

Muñoz-Jaramillo, Andrés (ORCID:0000000247160840)↗

The SENSEI Generic In Situ Interface: Tool and Processing Portability at Scale [Book Chapter]

One key challenge when doing in situ processing is the investment required to add code to numerical simulations needed to take advantage of in situ processing. Such instrumentation code is often specialized, and tailored to a specific in situ method or infrastructure. Then, if a simulation wants to use other in situ tools, each of which has its own bespoke API [4], then the simulation code team will quickly become overwhelmed with having a different set of instrumentation APIs, one per in situ tool or method. In an ideal situation, such instrumentation need happen only once, and then the instrumentation API provides access to a large diversity of tools. In this way, a data producer’s instrumentation need not be modified if the user desires to take advantage of a different set of in situ tools. The SENSEI generic in situ interface addresses this challenge, which means that SENSEI-instrumented codes enjoy the benefit of being able to use a diversity of tools at scale, tools that include Libsim, Catalyst, Ascent, as well as user-defined methods written in C++ or Python. SENSEI has been shown to scale to greater than 1M-way concurrency on HPC platforms, and provides support for a rich and diverse collection of common scientific data models. Furthermore, this chapter presents the key design challenges that enable tool and processing portability at scale, some performance analysis, and example science applications of the methods.

Bethel, E. Wes↗