Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

GTC Follow-up Observations of Very Metal-poor Star Candidates from DESI

Abstract The observations from the Dark Energy Spectroscopic Instrument (DESI) will significantly increase the numbers of known extremely metal-poor stars by a factor of ∼10, improving the sample statistics to study the early chemical evolution of the Milky Way and the nature of the first stars. In this paper we report follow-up observations with high signal-to-noise ratio of nine metal-poor stars identified during the DESI commissioning with the Optical System for Imaging and Low-Resolution Integrated Spectroscopy (OSIRIS) instrument on the 10.4 m Gran Telescopio Canarias. The analysis of the data using a well-vetted methodology confirms the quality of the DESI spectra and the performance of the pipelines developed for the data reduction and analysis of DESI data.

79 ASTRONOMY AND ASTROPHYSICS↗

Maintaining Trust in Reduction: Preserving the Accuracy of Quantities of Interest for Lossy Compression

As the growth of data sizes continues to outpace computational resources, there is a pressing need for data reduction techniques that can significantly reduce the amount of data and quantify the error incurred in compression. Compressing scientific data presents many challenges for reduction techniques since it is often on non-uniform or unstructured meshes, is from a high-dimensional space, and has many Quantities of Interests (QoIs) that need to be preserved. To illustrate these challenges, we focus on data from a large scale fusion code, XGC. XGC uses a Particle-In-Cell (PIC) technique which generates hundreds of PetaBytes (PBs) of data a day, from thousands of timesteps. XGC uses an unstructured mesh, and needs to compute many QoIs from the raw data, f.One critical aspect of the reduction is that we need to ensure that QoIs derived from the data (density, temperature, flux surface averaged momentums, etc.) maintain a relative high accuracy. We show that by compressing XGC data on the high-dimensional, nonuniform grid on which the data is defined, and adaptively quantizing the decomposed coefficients based on the characteristics of the QoIs, the compression ratios at various error tolerances obtained using a multilevel compressor (MGARD) increases more than ten times. We then present how to mathematically guarantee that the accuracy of the QoIs computed from the reduced f is preserved during the compression. We show that the error in the XGC density can be kept under a user-specified tolerance over 1000 timesteps of simulation using the mathematical QoI error control theory of MGARD, whereas traditional error control on the data to be reduced does not guarantee the accuracy of the QoIs.

Gong, Qian↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence↗

Cleaning Images with Gaussian Process Regression

Many approaches to astronomical data reduction and analysis cannot tolerate missing data: corrupted pixels must first have their values imputed. This paper presents astrofix, a robust and flexible image imputation algorithm based on Gaussian process regression. Through an optimization process, astrofix chooses and applies a different interpolation kernel to each image, using a training set extracted automatically from that image. It naturally handles clusters of bad pixels and image edges and adapts to various instruments and image types. For bright pixels, the mean absolute error of astrofix is several times smaller than that of median replacement and interpolation by a Gaussian kernel. We demonstrate good performance on both imaging and spectroscopic data, including the SBIG 6303 0.4 m telescope and the FLOYDS spectrograph of Las Cumbres Observatory and the CHARIS integral-field spectrograph on the Subaru Telescope.

42 ENGINEERING↗

Scalable Incremental Checkpointing using GPU-Accelerated De-Duplication

Writing large amounts of data concurrently to stable storage is a typical I/O pattern of many HPC workflows. This pattern introduces high I/O overheads and results in increased storage space utilization especially for workflows that need to capture the evolution of data structures with high frequency as checkpoints. In this context, many applications, such as graph pattern matching, perform sparse updates to large data structures between checkpoints. For these applications, incremental checkpointing techniques that save only the differences from one checkpoint to another can dramatically reduce the checkpoint sizes, I/O bottlenecks, and storage space utilization. However, such techniques are not without challenges: it is non-trivial to transparently determine what data has changed since a previous checkpoint and assemble the differences in a compact fashion that does not result in excessive metadata. State-of-art data reduction techniques (e.g., compression and de-duplication) have significant limitations when applied to modern HPC applications that leverage GPUs: slow at detecting the differences, generate a large amount of metadata to keep track of the differences, and ignore crucial spatiotemporal checkpoint data redundancy. This paper addresses these challenges by proposing a Merkle tree-based incremental checkpointing method to exploit GPUs' high memory bandwidth and massive parallelism. Experimental results at scale show a significant reduction of the I/O overhead and space utilization of checkpointing compared with state-of-the-art incremental checkpointing and compression techniques.

Tan, Nigel↗

Codebase release 2.0 for sauce

Low energy nuclear physics experiments are transitioning towards fully digital data acquisition systems. Realizing the gains in flexibility afforded by these systems relies on equally flexible data reduction techniques. In this paper, methods utilizing data frames and in-memory techniques to work with data, including data from self-triggering, digital data acquisition systems, are discussed within the context of a Python package, sauce. It is shown that data frame operations can encompass common analysis needs and allow interactive data analysis. Two event building techniques, dubbed referenced and referenceless event building, are shown to provide a means to transform raw list mode data into correlated multi-detector events. These techniques are demonstrated in the analysis of two example data sets.

Marshall, Caleb (ORCID:0000000211942920)↗

The ECP ALPINE project: In situ and post hoc visualization infrastructure and analysis capabilities for exascale

A significant challenge on an exascale computer is the speed at which we compute results exceeds by many orders of magnitude the speed at which we save these results. Therefore the Exascale Computing Project (ECP) ALPINE project focuses on providing exascale-ready visualization solutions including in situ processing. In situ visualization and analysis runs as the simulation is run, on simulations results are they are generated avoiding the need to save entire simulations to storage for later analysis. The ALPINE project made post hoc visualization tools, ParaView and VisIt, exascale ready and developed in situ algorithms and infrastructures. The suite of ALPINE algorithms developed under ECP includes novel approaches to enable automated data analysis and visualization to focus on the most important aspects of the simulation. Many of the algorithms also provide data reduction benefits to meet the I/O challenges at exascale. ALPINE developed a new lightweight in situ infrastructure, Ascent.

97 MATHEMATICS AND COMPUTING↗

The continuous readout stream of the MicroBooNE liquid argon time projection chamber for detection of supernova burst neutrinos

The MicroBooNE continuous readout stream is a parallel readout of the MicroBooNE liquid argon time projection chamber (LArTPC) which enables detection of non-beam events such as those from a supernova neutrino burst. The low energies of the supernova neutrinos and the intense cosmic-ray background flux due to the near-surface detector location makes triggering on these events very challenging. Instead, MicroBooNE relies on a delayed trigger generated by SNEWS (the Supernova Early Warning System) for detecting supernova neutrinos. The continuous readout of the LArTPC generates large data volumes, and requires the use of real-time compression algorithms (zero suppression and Huffman compression) implemented in an FPGA (field-programmable gate array) in the readout electronics. In this paper we present the results of the optimization of the data reduction algorithms, and their operational performance. To demonstrate the capability of the continuous stream to detect low-energy electrons, a sample of Michel electrons from stopping cosmic-ray muons is reconstructed and compared to a similar sample from the lossless triggered readout stream.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Prospects for Galactic and stellar astrophysics with asteroseismology of giant stars in the TESS continuous viewing zones and beyond

ABSTRACT The NASA Transiting Exoplanet Survey Satellite (NASA-TESS) mission presents a treasure trove for understanding the stars it observes and the Milky Way, in which they reside. We present a first look at the prospects for Galactic and stellar astrophysics by performing initial asteroseismic analyses of bright (G < 11) red giant stars in the TESS southern continuous viewing zone (SCVZ). Using three independent pipelines, we detect νmax and Δν in 41 per cent of the 15 405 star parent sample (6388 stars), with consistency at a level of $\sim \! 2{{\ \rm per\ cent}}$ in νmax and $\sim \! 5{{\ \rm per\ cent}}$ in Δν. Based on this, we predict that seismology will be attainable for ∼3 × 105 giants across the whole sky and at least 104 giants with ≥1 yr of observations in the TESS-CVZs, subject to improvements in analysis and data reduction techniques. The best quality TESS-CVZ data, for 5574 stars where pipelines returned consistent results, provide high-quality power spectra across a number of stellar evolutionary states. This makes possible studies of, for example, the asymptotic giant branch bump. Furthermore, we demonstrate that mixed ℓ = 1 modes and rotational splitting are cleanly observed in the 1-yr data set. By combining TESS-CVZ data with TESS-HERMES, SkyMapper, APOGEE, and Gaia, we demonstrate its strong potential for Galactic archaeology studies, providing good age precision and accuracy that reproduces well the age of high [α/Fe] stars and relationships between mass and kinematics from previous studies based on e.g. Kepler. Better quality astrometry and simpler target selection than the Kepler sample makes this data ideal for studies of the local star formation history and evolution of the Galactic disc. These results provide a strong case for detailed spectroscopic follow-up in the CVZs to complement that which has been (or will be) collected by current surveys.

Mackereth, J. Ted↗

Information-Theoretic Exploration of Multivariate Time-Varying Image Databases

Modern scientific simulations produce very large datasets, making interactive exploration of such data computationally prohibitive. An increasingly common data reduction technique is to store visualizations and other data extracts in a database. The Cinema project is one such approach, storing visualizations in an image database for post hoc exploration and interactive image-based analysis. This work focuses on developing efficient algorithms that can quantify various types of multivariate dependencies existing within multi-variable datasets. It applies specific mutual information measures for the quantification of salient regions from multivariate image data. Here, using such information measures, the opacity of the images is modulated so that the salient regions are automatically highlighted and the domain scientists can interactively explore the most relevant regions for scientific discovery.

97 MATHEMATICS AND COMPUTING↗

An algorithm for resolving intragranular orientation fields using coupled far-field and near-field high energy $\mathrm{X}$-ray diffraction microscopy

Here, we present a novel algorithm for reconstructing spatial intragranular lattice orientation fields using both far-field and near-field high energy X-ray diffraction microscopy (HEDM) measurements. An established far-field indexing algorithm is modified to include lattice orientation distribution information (grain orientation envelopes) in addition to average grain orientations. The near-field data reduction algorithm utilizes this enriched far-field orientation data as a seed for reconstructing intragranular spatial maps of lattice orientation from the diffraction images. The primary benefit of the new algorithm is a significant decrease in the number of trial calculations that must be performed in the reconstruction process compared to existing methodologies while maintaining scalability. The resulting gains in efficiency facilitate the use of relatively modest computational resources and improve throughput at the point of measurement. We provide two example applications: volumetric orientation field reconstructions for a Ti-Al alloy both before and after the application of 3% uniaxial strain. The results showcase the efficiency of the new method and the ability to resolve subtle changes in microstructure, which are associated with incipient plastic deformation.

36 MATERIALS SCIENCE↗

A Co-design Framework for Online Data Analysis and Reduction

Science applications preparing for the exascale era are increasingly exploring in situ computations comprising of simulation-analysis-reduction pipelines coupled in-memory. Efficient composition and execution of such complex pipelines for a target platform is a codesign process that evaluates the impact and tradeoffs of various application- and system-specific parameters. In this article, we describe a toolset for automating performance studies of composed HPC applications that perform online data reduction and analysis. We describe Cheetah, a new framework for composing parametric studies on coupled applications, and Savanna, a runtime engine for orchestrating and executing campaigns of codesign experiments. Furthermore, this toolset facilitates understanding the impact of various factors such as process placement, synchronicity of algorithms, and storage versus compute requirements for online analysis of large data. Ultimately, we aim to create a catalog of performance results that can help scientists understand tradeoffs when designing next-generation simulations that make use of online processing techniques. We illustrate the design of Cheetah and Savanna, and present application examples that use this framework to conduct codesign studies on small clusters as well as leadership class supercomputers.

97 MATHEMATICS AND COMPUTING↗

Multiresolution classification of turbulence features in image data through machine learning

During large-scale simulations, intermediate data products such as image databases have become popular due to their low relative storage cost and fast in-situ analysis. Serving as a form of data reduction, these image databases have become more acceptable to perform data analysis on. In this work, we present an image-space detection and classification system for extracting vortices at multiple scales through wavelet-based filtering. A custom image-space descriptor is used to encode a large variety of vortex-types and a machine learning system is trained for fast classification of vortex regions. By combining a radial-based histogram descriptor, a bag of visual words feature descriptor, and a support vector machine, our results show that we are able to detect and classify vortex features at various sizes at multiple scales. Once trained, our framework enables the fast extraction of vortices on new, unknown image datasets for flow analysis.

97 MATHEMATICS AND COMPUTING↗

Workflows Community Summit 2022: A Roadmap Revolution

Scientific workflows have become integral tools in broad scientific computing use cases. Science discovery is increasingly dependent on workflows to orchestrate large and complex scientific experiments that range from the execution of a cloud-based data preprocessing pipeline to multi-facility instrument-to-edge-to-HPC computational workflows. Given the changing landscape of scientific computing (often referred to as a computing continuum) and the evolving needs of emerging scientific applications, it is paramount that the development of novel scientific workflows and system functionalities seek to increase the efficiency, resilience, and pervasiveness of existing systems and applications. Specifically, the proliferation of machine learning/artificial intelligence (ML/AI) workflows, need for processing large-scale datasets produced by instruments at the edge, intensification of near real-time data processing, support for long-term experiment campaigns, and emergence of quantum computing as an adjunct to HPC, have significantly changed the functional and operational requirements of workflow systems. Workflow systems now need to, for example, support data streams from the edge-to-cloud-to-HPC, enable the management of many small-sized files, allow data reduction while ensuring high accuracy, orchestrate distributed services (workflows, instruments, data movement, provenance, publication, etc.) across computing and user facilities, among others. Further, to accelerate science, it is also necessary that these systems implement specifications/standards and APIs for seamless (horizontal and vertical) integration between systems and applications, as well as enable the publication of workflows and their associated products according to the FAIR principles.

97 MATHEMATICS AND COMPUTING↗

First M87 Event Horizon Telescope Results. VII. Polarization of the Ring

In 2017 April, the Event Horizon Telescope (EHT) observed the near-horizon region around the supermassive black hole at the core of the M87 galaxy. These 1.3 mm wavelength observations revealed a compact asymmetric ring-like source morphology. This structure originates from synchrotron emission produced by relativistic plasma located in the immediate vicinity of the black hole. Here we present the corresponding linear-polarimetric EHT images of the center of M87. We find that only a part of the ring is significantly polarized. The resolved fractional linear polarization has a maximum located in the southwest part of the ring, where it rises to the level of ~15%. The polarization position angles are arranged in a nearly azimuthal pattern. We perform quantitative measurements of relevant polarimetric properties of the compact emission and find evidence for the temporal evolution of the polarized source structure over one week of EHT observations. The details of the polarimetric data reduction and calibration methodology are provided. We carry out the data analysis using multiple independent imaging and modeling techniques, each of which is validated against a suite of synthetic data sets. The gross polarimetric structure and its apparent evolution with time are insensitive to the method used to reconstruct the image. These polarimetric images carry information about the structure of the magnetic fields responsible for the synchrotron emission. Their physical interpretation is discussed in an accompanying publication.

79 ASTRONOMY AND ASTROPHYSICS↗

A Digital Twin Framework Utilizing Machine Learning for Robust Predictive Maintenance: Enhancing Tire Health Monitoring

We introduce a novel digital twin (DT) framework for the predictive maintenance of long-term physical systems. Using monitoring tire health as an application, we show how the DT framework can be used to enhance automotive safety and efficiency, and how the technical challenges can be overcome using a three-step approach. First, to manage the data complexity over a long operation span, we employ data reduction techniques to concisely represent physical tires using historical performance and usage data. Relying on these data, for fast real-time prediction, we train a transformer-based model offline on our concise dataset to predict future tire health over time, represented as remaining casing potential (RCP). Based on our architecture, our model quantifies both epistemic and aleatoric uncertainties, providing reliable confidence intervals around predicted RCP. Second, to incorporate real-time data, we update the predictive model in the DT framework, ensuring its accuracy throughout its lifespan with the aid of hybrid modeling and the use of the discrepancy function. Third, to assist decision-making in predictive maintenance, we implement a tire state decision algorithm, which strategically determines the optimal timing for tire replacement based on RCP forecasted by our transformer model. This approach ensures that our DT accurately predicts system health, continually refines its digital representation, and supports predictive maintenance decisions. Furthermore, our framework effectively embodies a physical system, leveraging big data and machine learning (ML) for predictive maintenance, model updates, and decision-making.

advanced computing infrastructure↗

Advanced Image Reconstruction for MCP Detector in Event Mode

A two-step data reduction framework is proposed in this study to reconstruct a radiograph from the data collected with a micro-channel plate (MCP) detector operating under event mode. One clustering algorithm and three neutron event back-tracing models are proposed and evaluated using both example data and a full scan data. The reconstructed radiographs are analyzed, the results of which are used to suggest future development.

Zhang, Chen↗