Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Bringing Different Views Together: A Hybrid Cooperative Perception Framework for Connected Autonomous Vehicles

Cooperative perception will be essential for connected autonomous vehicles to enhance object recognition and optimize path planning by sending data information about the surrounding environment. However, an inherent challenge in existing systems is the high bandwidth cost of transmitting information in real-time, which restricts cooperative perception’s practicality. Here, this work presents a hybrid cooperative perception fusion framework aimed at mitigating this issue by optimizing data transmission according to available bandwidth or through data reduction techniques. Our methods ensure that vehicles can rapidly transmit high-confidence data without overwhelming the network. Experimental results indicate that our methodology substantially diminishes data transmission sizes while maintaining object detection accuracy. For cooperative perception in autonomous vehicle systems, our approach provides a scalable and effective way to get past the bandwidth barrier.

Carrillo, Dominic [Univ. of North Texas, Denton, T↗

Progressive Tree-Based Compression of Large-Scale Particle Data

Scientific simulations and observations using particles have been creating large datasets that require effective and efficient data reduction to store, transfer, and analyze. However, current approaches either compress only small data well while being inefficient for large data, or handle large data but with insufficient compression. Toward effective and scalable compression/decompression of particle positions, we introduce new kinds of particle hierarchies and corresponding traversal orders that quickly reduce reconstruction error while being fast and low in memory footprint. Our solution to compression of large-scale particle data is a flexible block-based hierarchy that supports progressive, random-access, and error-driven decoding, where error estimation heuristics can be supplied by the user. For low-level node encoding, we introduce new schemes that effectively compress both uniform and densely structured particle distributions. Our proposed methods thus target all three phases of a tree-based particle compression pipeline, namely tree construction, tree traversal, and node encoding. In conclusion, the improved efficacy and flexibility of these methods over existing compressors are demonstrated through extensive experimentation, using a wide range of scientific particle datasets.

97 MATHEMATICS AND COMPUTING↗

Design and Testing of the Endcap Concentrator ASICs for the CMS High-Granularity Calorimeter Upgrade

A major upgrade of the High-Granularity Calorimeter (HGCAL) in the CMS detector is planned for Long-Shutdown 3 (LS3), currently expected to start in 2026. This upgrade will contain over 6 million channels and the electronics will be required to be low power and to withstand a radiation environment with a High-Energy Hadron flux of 3x10**6 particles per square centimeter per second. The solution to this significant challenge is the two Endcap Concentrator (ECON) ASICs, ECON-T and ECON-D, working in tandem with the HGCROC front-end ASIC. The two ECON ASICs provide critical on-detector data reduction for both the 40 MHz trigger path (ECON-T) and 750 kHz data acquisition path (ECON-D) of the HGCAL. The ASICs are fabricated in 65nm CMOS. They are rad-tolerant to 600 Mrad with low power consumption (<2.5 mW/channel). This presentation will be a comprehensive description of each ECON design, including the infrastructure that they share as well as the elements that make each unique. The presentation will also include functionality and radiation tests for both ASICs, and the first high statistics characterization results from the full production of 75k ECON-D and ECON-T ASICs.

Hoff, James R. [Fermilab] (ORCID:0000000163514592)↗

The LSST DESC data challenge 1: generation and analysis of synthetic images for next-generation surveys

Data Challenge 1 (DC1) is the first synthetic data set produced by the Rubin Observatory Legacy Survey of Space and Time (LSST) Dark Energy Science Collaboration (DESC). DC1 is designed to develop and validate data reduction and analysis and to study the impact of systematic effects that will affect the LSST data set. DC1 is comprised of r -band observations of 40 deg 2 to 10 yr LSST depth. In this paper, we present each stage of the simulation and analysis process: (a) generation, by synthesizing sources from cosmological N -body simulations in individual sensor-visit images with different observing conditions; (b) reduction using a development version of the LSST Science Pipelines; and (c) matching to the input cosmological catalogue for validation and testing. We verify that testable LSST requirements pass within the fidelity of DC1. We establish a selection procedure that produces a sufficiently clean extragalactic sample for clustering analyses and we discuss residual sample contamination, including contributions from inefficiency in star–galaxy separation and imperfect deblending. We compute the galaxy power spectrum on the simulated field and conclude that: (i) survey properties have an impact of 50 per cent of the statistical uncertainty for the scales and models used in DC1; (ii) a selection to eliminate artefacts in the catalogues is necessary to avoid biases in the measured clustering; and (iii) the presence of bright objects has a significant impact (2σ–6σ) in the estimated power spectra at small scales (ℓ > 1200), highlighting the impact of blending in studies at small angular scales in LSST.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)↗

Fast Semi-automated Filtration Method for Non-targeted LC-QTOF Data of Aged Nitroplasticizer Samples

A full dataset of aged nitroplasticizer (NP) is composed of more than 2000 unique mass-to-charges (m/z) when combining the non-targeted data obtained from both positive and negative electrospray ionization modes in time-of-flight mass spectrometry. Therefore, manual processing of these data often takes days, weeks, or even months to scrutinize for mechanistic insights. To effectively extract meaningful signals that represent vital degradation intermediates in the early NP degradation mechanism, a semi-automated postprocessing workflow for data filtering, tailored to the aging experiment of NP, has been developed. The automated portion of this workflow is written in a Python code (using pandas, numpy, and matplotlib libraries), which removes more than 65% of potential false signals within seconds via four threshold-based adjustable filters: signal sensitivity, coefficient of variation, number of measurements, and retention time variability. As for the manual portion, a pattern-based inspection method is employed to reduce another 23% or more false positives, which greatly simplifies data visualization and results in less than 3% of potential candidate m/z needing in-depth data interpretation. As a positive control, known compounds are verified. Using this semi-automated data reduction method, the amount of time required is reduced to a matter of hours for data filtering in the non-targeted datasets of aged NP, which saves more time and effort for compound identification.

36 MATERIALS SCIENCE↗

Experimental methods for laboratory measurements of helium spectral line broadening in white dwarf photospheres

White Dwarf (WD) stars are the most common stellar remnant in the universe. WDs usually have a hydrogen or helium atmosphere, and helium WD (called DB) spectra can be used to solve outstanding problems in stellar and galactic evolution. DB origins, which are still a mystery, must be known to solve these problems. DB masses are crucial for discriminating between different proposed DB evolutionary hypotheses. Current DB mass determination methods deliver conflicting results. The spectroscopic mass determination method relies on line broadening models that have not been validated at DB atmosphere conditions. We performed helium benchmark experiments using the White Dwarf Photosphere Experiment (WDPE) platform at Sandia National Laboratories' Z-machine that aims to study He line broadening at DB conditions. Using hydrogen/helium mixture plasmas allows investigating the importance of He Stark and van der Waals broadening simultaneously. Accurate experimental data reduction methods are essential to test these line-broadening theories. In this paper, we present data calibration methods for these benchmark He line shape experiments. We give a detailed account of data processing, spectral power calibrations, and instrument broadening measurements. Uncertainties for each data calibration step are also derived. We demonstrate that our experiments meet all benchmark experiment accuracy requirements: WDPE wavelength uncertainties are <1 Å, spectral powers can be determined to within 15%, densities are accurate at the 20% level, and instrumental broadening can be measured with 20% accuracy. Fulfilling these stringent requirements enables WDPE experimental data to provide physically meaningful conclusions about line broadening at DB conditions.

79 ASTRONOMY AND ASTROPHYSICS↗

Efficient Data Query for Gaussian Process Compressed Data through Value Range Estimation [Slides]

When the resolution of the data increases, data reduction methods are applied to simulation output, including Gaussian process, neural representation and compression algorithms. Lots of data analysis/visualization techniques requires data query, but data query from reduced representation is still challenging. This report will provide examples and provide possible answers to why data query from reduced representation is still challenging.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Archetype-based Redshift Estimation for the Dark Energy Spectroscopic Instrument Survey

We present a computationally efficient galaxy archetype-based redshift estimation and spectral classification method for the Dark Energy Survey Instrument (DESI) survey. The DESI survey currently relies on a redshift fitter and spectral classifier using a linear combination of principal component analysis–derived templates, which is very efficient in processing large volumes of DESI spectra within a short time frame. However, this method occasionally yields unphysical model fits for galaxies and fails to adequately absorb calibration errors that may still be occasionally visible in the reduced spectra. Our proposed approach improves upon this existing method by refitting the spectra with carefully generated physical galaxy archetypes combined with additional terms designed to absorb data reduction defects and provide more physical models to the DESI spectra. We test our method on an extensive data set derived from the survey validation (SV) and Year 1 (Y1) data of DESI. Our findings indicate that the new method delivers marginally better redshift success for SV tiles while reducing catastrophic redshift failure by 10%–30%. At the same time, results from millions of targets from the main survey show that our model has relatively higher redshift success and purity rates (0.5%–0.8% higher) for galaxy targets while having similar success for QSOs. These improvements also demonstrate that the main DESI redshift pipeline is generally robust. Additionally, it reduces the false-positive redshift estimation by 5%–40% for sky fibers. We also discuss the generic nature of our method and how it can be extended to other large spectroscopic surveys, along with possible future improvements.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Film condensation with high heat fluxes and scaled experiments using pure steam for reactor containment cooling

Condensation tests were performed using a newly developed test facility for scaling the passive containment cooling system (PCCS) to a small modular reactor (SMR). The PCCS of the SMR plays a pivotal role in ensuring greater safety, reliability, and compactness than what is afforded by traditional reactors. Therefore, a well-designed PCCS is essential to SMRs. However, previous studies and test data were unsuitable for scaling, due to high variation in the test geometry and operating conditions. This study intends to close this research gap by using a novel designed scaled test facility consisting of vertical condensing test sections featuring 1-, 2-, and 4-inch-diameter condensing tubes with annular water cooling, and by applying superheated and saturated steam with different steam mass flow ranges of 5–25 g/s. Further, the primary test data, including axial temperatures, mass flow rates, and pressures, were used in conjunction with a standard data reduction method to estimate critical parameters such as heat fluxes, heat transfer coefficients, and condensation rates. These scaled test data would support improving empirical correlations and validating condensation models to identify scaling distortion for SMR PCCSs.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Big Data Meets Geothermal Exploration (CRADA Final Report)

As part of the Cyclotron Road program, Zanskar Geothermal & Minerals, Inc. investigated the application of micro-earthquake and ambient noise seismology methods to imaging and characterizing the structural characteristics and hydrothermal flux of subsurface faults. Significant advances in what could be resolved were enabled by two major developments in seismology: 1) the availability of large-n arrays of low-cost seismometers, and 2) the availability of increased computational power and semi-automated data reduction algorithms. In tandem, these advances may improve the signal-to-noise ratio and spatial precision of the data collected and enable higher-resolution characterization of subsurface fracture systems and their spatio-temporal evolution. These tools supported efforts to reduce dry-hole risk and to improve wellfield productivity for geothermal resource development. In particular, two applications of these advances were evaluated: 1) fracture-seismic imaging, which was used to detect ambient emissions from fluid-filled fractures, and 2) reservoir tomography, which used information about travel paths, source locations, and source parameters of micro-earthquakes to identify areas of enhanced permeability. Integration of these methods provided guidance for siting wells and served as prior constraints for reservoir models, informing forecasts of power potential and production and injection strategies aimed at minimizing temperature decline and improving overall resource productivity.

15 GEOTHERMAL ENERGY↗

TensorID v1.0

This Python software package includes new and efficient algorithms for satellite and core interpolative decomposition of tensor data. In general, these algorithms target high-dimensional data reduction and compression. The software is purely numerical and can be applied by others to many important sources of tensor data generated by computation or experiment.

Zhang, Yifan [Lawrence Berkeley National Laborator↗

Towards Autonomous Experiments by Connecting High Performance Microscopy with High Performance Computing

The digitization of controls, data, and analysis in microscopy is bringing the idea of autonomous microscopes closer to reality than ever before. Automated transmission electron microscopy (TEM) is already fairly routine for some experiments the only require simple repetitive tasks such as imaging biological macromolecules for single particle cryoEM [1], tilt series for electron tomography [2], and movies for crystallography [3]. The vast majority of TEM experiments are conducted completely by human operators who choose the regions of interest, optimize experimental parameters, and make decisions about data quality visually during an experiment. The field is still a long way from having completely autonomous TEMs that can adapt to sample difficulties and tune experimental parameters based on data quality and desired experimental outcomes. Part of the issue is the lack of capability for feeding information learned from on-line, live data analysis back into the on-going experiment [4]. Furthermore, this presentation will discuss current capabilities for large scale data reduction and analysis using high performance computing (i.e. supercomputing) and progress towards developing a true feed-back loop that places data analysis and theory in the experimental loop.

97 MATHEMATICS AND COMPUTING↗

Snowmass Letter of Interest - Analysis Facilities - CompF5

In this letter, we comment on the “last mile” for computing in support of physics analyses; this does not cover simulation/acquisition, batch processing/production/reconstruction, and data reduction. Specifically, we focus on analysis facilities, which provide the foundation for the individual researcher to experiment with and understand their data.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Methods for Incorporating Model Uncertainty into Exoplanet Atmospheric Analysis

A key goal of exoplanet spectroscopy is to measure atmospheric properties, such as abundances of chemical species, in order to connect them to our understanding of atmospheric physics and planet formation. In this new era of high-quality JWST data, it is paramount that these measurement methods are robust. When comparing atmospheric models to observations, multiple candidate models may produce reasonable fits to the data. Typically, conclusions are reached by selecting the best-performing model according to some metric. This ignores model uncertainty in favor of specific model assumptions, potentially leading to measured atmospheric properties that are overconfident and/or incorrect. In this paper, we compare three ensemble methods for addressing model uncertainty by combining posterior distributions from multiple analyses: Bayesian model averaging, a variant of Bayesian model averaging using leave-one-out predictive densities, and stacking of predictive distributions. We demonstrate these methods by fitting the Hubble Space Telescope (HST) + Spitzer transmission spectrum of the hot Jupiter HD 209458b using models with different cloud and haze prescriptions. All of our ensemble methods lead to uncertainties on retrieved parameters that are larger but more realistic and consistent with physical and chemical expectations. Since they have not typically accounted for model uncertainty, uncertainties of retrieved parameters from HST spectra have likely been underreported. We recommend stacking as the most robust model combination method. Our methods can be used to combine results from independent retrieval codes and from different models within one code. They are also widely applicable to other exoplanet analysis processes, such as combining results from different data reductions.

79 ASTRONOMY AND ASTROPHYSICS↗

POWTEX visits POWGEN

The high-intensity time-of-flight (TOF) neutron diffractometer POWTEX for powder and texture analysis is currently being built prior to operation in the eastern guide hall of the research reactor FRM II at Garching close to Munich, Germany. Because of the world-wide 3 He crisis in 2009, the authors promptly initiated the development of 3 He-free detector alternatives that are tailor-made for the requirements of large-area diffractometers. Herein is reported the 2017 enterprise to operate one mounting unit of the final POWTEX detector on the neutron powder diffractometer POWGEN at the Spallation Neutron Source located at Oak Ridge National Laboratory, USA. As a result, presented here are the first angular- and wavelength-dependent data from the POWTEX detector, unfortunately damaged by a 50 g shock but still operating, as well as the efforts made both to characterize the transport damage and to successfully recalibrate the voxel positions in order to yield nonetheless reliable measurements. Also described is the current data reduction process using the PowderReduceP2D algorithm implemented in Mantid [Arnold et al. (2014). Nucl. Instrum. Methods Phys. Res. A , 764 , 156–166]. The final part of the data treatment chain, namely a novel multi-dimensional refinement using a modified version of the GSAS-II software suite [Toby & Von Dreele (2013). J. Appl. Cryst. 46 , 544–549], is compared with a standard data treatment of the same event data conventionally reduced as TOF diffraction patterns and refined with the unmodified version of GSAS-II . This involves both determining the instrumental resolution parameters using POWGEN's powdered diamond standard sample and the refinement of a friendly-user sample, BaZn(NCN) 2 . Although each structural parameter on its own looks similar upon comparing the conventional (1D) and multi-dimensional (2D) treatments, also in terms of precision, a closer view shows small but possibly significant differences. For example, the somewhat suspicious proximity of the a and b lattice parameters of BaZn(NCN) 2 crystallizing in Pbca as resulting from the 1D refinement (0.008 Å) is five times less pronounced in the 2D refinement (0.038 Å). Similar features are found when comparing bond lengths and bond angles, e.g. the two N—C—N units are less differently bent in the 1D results (173 and 175°) than in the 2D results (167 and 173°). The results are of importance not only for POWTEX but also for other neutron TOF diffractometers with large-area detectors, like POWGEN at the SNS or the future DREAM beamline at the European Spallation Source.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗