Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

An Insight-Centric Paradigm for Data Reduction and Inference Speed Improvement at the Scurry Area Canyon Reef Operator’s Committee (SACROC) Unit

The poster presents the work conducted under SMART focusing on using insight-centric approach to design a meaningful proxy for machine learning. Domain insights are critical not just in understanding the prediction results but also in designing the model. This study demonstrated that a single meaningful scaler (as an extreme case) can effectively replace full-size 3D geologic properties. The model's accuracies are on par with other models, and it is the fastest model to predict all test cases, 5,000 times faster than traditional simulations.

Shih, Chung Yan↗

Towards Resilient Near Real-Time Analysis Workflows in Fusion Energy Science

Nuclear fusion holds the promise of an endless source of energy. Several research experiments across the world and joint modeling and simulation efforts between the nuclear physics and high performance computing communities are actively preparing the operation of the International Thermonuclear Experimental Reactor (ITER). Both experimental reactors and their simulated counterparts generate data that must be analyzed quickly and in a resilient way to support decision making for the configuration of subsequent runs or prevent a catastrophic failure. However, the cost if the traditional techniques used to improve the resilience of analysis workflows, i.e., replicating datasets and computational tasks, becomes prohibitive with explosion of the volume of data produced by modern instruments and simulations. Therefore, we advocate in this paper for an alternate approach based on data reduction and data streaming. The rationale is that by allowing for a reasonable, controlled, and guaranteed loss of accuracy it becomes possible to transfer smaller amounts of data, shorten the execution time of analysis workflows, and lower the cost of replication to increase resilience. We develop our research and development roadmap towards resilient near real-time analysis workflows in fusion energy science and present early results showing that data streaming and data reduction is a promising way to speed up the execution and improve the resilience of analysis workflows.

Suter, Fred↗

Multimodal X-ray nano-spectromicroscopy analysis of chemically heterogeneous systems

Abstract Understanding the nanoscale chemical speciation of heterogeneous systems in their native environment is critical for several disciplines such as life and environmental sciences, biogeochemistry, and materials science. Synchrotron-based X-ray spectromicroscopy tools are widely used to understand the chemistry and morphology of complex material systems owing to their high penetration depth and sensitivity. The multidimensional (4D+) structure of spectromicroscopy data poses visualization and data-reduction challenges. This paper reports the strategies for the visualization and analysis of spectromicroscopy data. We created a new graphical user interface and data analysis platform named XMIDAS (X-ray multimodal image data analysis software) to visualize spectromicroscopy data from both image and spectrum representations. The interactive data analysis toolkit combined conventional analysis methods with well-established machine learning classification algorithms (e.g. nonnegative matrix factorization) for data reduction. The data visualization and analysis methodologies were then defined and optimized using a model particle aggregate with known chemical composition. Nanoprobe-based X-ray fluorescence (nano-XRF) and X-ray absorption near edge structure (nano-XANES) spectromicroscopy techniques were used to probe elemental and chemical state information of the aggregate sample. We illustrated the complete chemical speciation methodology of the model particle by using XMIDAS. Next, we demonstrated the application of this approach in detecting and characterizing nanoparticles associated with alveolar macrophages. Our multimodal approach combining nano-XRF, nano-XANES, and differential phase-contrast imaging efficiently visualizes the chemistry of localized nanostructure with the morphology. We believe that the optimized data-reduction strategies and tool development will facilitate the analysis of complex biological and environmental samples using X-ray spectromicroscopy techniques.

36 MATERIALS SCIENCE↗

AMM: Adaptive Multilinear Meshes

Adaptive representations are increasingly indispensable for reducing the in-memory and on-disk footprints of large-scale data. Usual solutions are designed broadly along two themes: reducing data precision, e.g., through compression, or adapting data resolution, e.g., using spatial hierarchies. Additionally, recent research suggests that combining the two approaches, i.e., adapting both resolution and precision simultaneously, can offer significant gains over using them individually. However, there currently exist no practical solutions to creating and evaluating such representations at scale. In this work, we present a new resolution-precision-adaptive representation to support hybrid data reduction schemes and offer an interface to existing tools and algorithms. Through novelties in spatial hierarchy, our representation, Adaptive Multilinear Meshes (AMM), provides considerable reduction in the mesh size. AMM creates a piecewise multilinear representation of uniformly sampled scalar data and can selectively relax or enforce constraints on conformity, continuity, and coverage, delivering a flexible adaptive representation. AMM also supports representing the function using mixed-precision values to further the achievable gains in data reduction. We describe a practical approach to creating AMM incrementally using arbitrary orderings of data and demonstrate AMM on six types of resolution and precision datastreams. By interfacing with state-of-the-art rendering tools through VTK, we demonstrate the practical and computational advantages of our representation for visualization techniques. With an open-source release of our tool to create AMM, we make such evaluation of data reduction accessible to the community, which we hope will foster new opportunities and future data reduction schemes.

97 MATHEMATICS AND COMPUTING↗

Machine-learning-assisted automation of single-crystal neutron diffraction

Neutron scattering is a powerful but expensive technique to study materials and discover new matter. Advanced detector technology has significantly improved the efficiency of neutron experiments, increasing the complexity of neutron data reduction and analysis. Machine learning (ML) brings new directions for neutron diffraction data reduction and experiment operation. Here, this work presents an ML-assisted data reduction and analysis method for precise recognition of Bragg peaks and the corresponding regions of interest; it can then automatically screen and align a measured crystal using the recognized peaks, and subsequently plan and optimize the data collection with user-provided information and uncertainty quantification values of detected peaks. This method shows robust performance in different complex sample environments and enables automated single-crystal neutron diffraction.

47 OTHER INSTRUMENTATION↗

Robustness of the smartpixels classifier for different simulated sensor geometries and non-ideal detector conditions

Pixel tracking detectors at upcoming collider experiments will see unprecedented charged-particle densities. Real-time data reduction on the detector will enable higher granularity and faster readout, possibly enabling the use of the pixel detector in high-rate online event selection, such as the ATLAS or CMS first-level trigger systems. This data reduction can be accomplished with a neural network (NN) in the readout chip bonded with the sensor that recognizes and rejects tracks with low transverse momentum (p T ) based on the geometrical shape of the charge deposition (“cluster”). To design viable detectors for deployment, the dependence of the NN as a function of the sensor geometry, external magnetic field, irradiation, and noise must be understood. In this paper, we present first studies of the efficiency and data reduction for planar pixel sensors exploring these parameters. For the CMS HL-LHC sensor geometry, we obtain a signal efficiency of (91.9 ± 0.7)% and a data reduction of (29.7 ± 1.0)%. A smaller sensor pitch in the bending direction improves the p T discrimination, but a larger pitch can be partially compensated with detector thickness. Any accumulated radiation damage also changes the cluster shape, reducing the signal efficiency compared to the baseline by approximately 30–60% in absolute terms, but nearly all of the performance can be recovered through retraining of the network and updating the weights. Finally, the impact of noise was investigated, and retraining the network on noise-injected datasets was found to maintain performance within 6% of the baseline network trained and evaluated on noiseless data. •ASIC-compatible track-momentum classifier is robust in realistic detector conditions.•About 90% signal efficiency and 30% data reduction per layer for CMS HL-LHC geometry.•Single-layer signal efficiency increases for smaller pixel pitch or thicker sensors.•Performance with noise or after radiation damage mostly recovered by retraining.

Shekar, Danush [Illinois U., Chicago] (ORCID:00000↗

Tools for Visualization and Analysis of Small-Angle Neutron Scattering Data: Descriptions and Examples

A great deal of progress has been made in improving the data reduction experience for the SANS instruments at the SNS and HFIR at ORNL. The existing data reduction toolset, drtsans, makes it possible to integrate data analysis and visualization tools into the data reduction scripts, thereby providing new opportunities for more automated data processing for users of the SNS and HFIR. Here, the first set of tools developed is described with usage examples.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

ECON-T and ECON-D: Endcap Concentrator ASICs for the CMS HGCAL

With over 6 million channels, the High Granularity Calorimeter (HGCAL) for the CMS HL-LHC Upgrade presents a unique data challenge. The ECON ASICs provide critical on-detector data reduction for the 40 MHz trigger path (ECON-T) and 750 kHz data acquisition path (ECON-D) of the HGCAL. The ASICs, fabricated in 65 nm CMOS, are rad-tolerant (600 Mrad) with low power consumption (<2.5 mW/channel). This presentation is the first comprehensive description of the ECON designs, first functionality and radiation tests for the ECON-T ASIC, and first results from the full production of 75k ECON-D and ECON-T ASICs.

Bergamin, G. [CERN] (ORCID:0000000285758704)↗

Online data analysis and reduction: An important co-design motif for extreme-scale computers

A growing disparity between supercomputer computation speeds and I/O rates means that it is rapidly becoming infeasible to analyze supercomputer application output only after that output has been written to a file system. Instead, data-generating applications must run concurrently with data reduction and/or analysis operations, with which they exchange information via high-speed methods such as interprocess communications. The resulting parallel computing motif, online data analysis and reduction (ODAR), has important implications for both application and HPC systems design. Here we introduce the ODAR motif and its co-design concerns, describe a co-design process for identifying and addressing those concerns, present tools that assist in the co-design process, and present case studies to illustrate the use of the process and tools in practical settings.

Data Analysis↗

Developing ML/AI Methods for High-Throughput Characterization of Multiple-Sensor Streams of Tokamak Dynamics for High-Speed Control (Final Report)

This project evaluated and developed new mathematical and algorithmic techniques capable of handling (in real-time) the growing amounts of data generated by modern fusion research. While existing numerical linear algebra (NLA) methods provide the backbone to classical data analysis and algorithms, these methods fundamentally do not port to distributed architectures nor do they allow low-latency data reduction for control. Motivated by the needs for modern fusion reactors, this project explored and implemented new numerical methods to characterize plasma dynamics, respond in real-time to discharge evolution, and to process massive-scale data accurately and rapidly more fully. This project links expertise in multiple-sensor diagnostics of tokamak plasma dynamics from Columbia University’s Plasma Physics Laboratory with expertise in massive-scale data reduction and extreme data control algorithms at Columbia University’s Data Science Institute. This interdisciplinary project (i) applied machine learning methods, (ii) implemented a properly-trained neural-network for very fast processing of high-speed plasma videography, and (ii) developed the applied mathematical methods, based on randomized-NLA (rNLA) routines, for data analysis, reduction, and real-time control. The Columbia University High Beta Tokamak-Extended Pulse (HBT-EP) facility provided data to test new algorithms and partnership with Columbia University's Data Sciences Institute evaluated the broader use of new algorithms for many challenging control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

User Guide for Sample Reduction at GP-SANS

This manual is intended as a quick guide for data reduction of GP-SANS data. It includes all necessary steps to do the data reduction based on absolute calibration using the open beam method and how to transfer the reduced data to the personal computer system. If any errors are coming up so that the reduction script is not functioning as intended, please contact the instrument scientist.

97 MATHEMATICS AND COMPUTING↗

Another Set of Python Tools for Visualizing and Manipulating Small-Angle Neutron Scattering Data: Descriptions and Examples

The GP-SANS, Bio-SANS and EQ-SANS instruments at ORNL utilize drtsans for data reduction. drtsans is built on Python, and it can be run using python scripts and Jupyter notebooks. The flexibility afforded by Python makes it possible to incorporate additional actions into the scripts used for data reduction, such as analysis and visualization. Here, a new set of tools for visualizing and manipulating SANS data that can be incorporated into the data reduction scripts for the ORNL SANS instruments, or employed during post–processing, is presented that expands the capabilities of the two previously-released tool sets.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Sensor Co-design for $\textit{smartpixels}$

Pixel tracking detectors at upcoming collider experiments will see unprecedented charged-particle densities. Real-time data reduction on the detector will enable higher granularity and faster readout, possibly enabling the use of the pixel detector in the first level of the trigger for a hadron collider. This data reduction can be accomplished with a neural network (NN) in the readout chip bonded with the sensor that recognizes and rejects tracks with low transverse momentum (p$_T$) based on the geometrical shape of the charge deposition (``cluster''). To design a viable detector for deployment at an experiment, the dependence of the NN as a function of the sensor geometry, external magnetic field, and irradiation must be understood. In this paper, we present first studies of the efficiency and data reduction for planar pixel sensors exploring these parameters. A smaller sensor pitch in the bending direction improves the p$_T$ discrimination, but a larger pitch can be partially compensated with detector depth. An external magnetic field parallel to the sensor plane induces Lorentz drift of the electron-hole pairs produced by the charged particle, broadening the cluster and improving the network performance. The absence of the external field diminishes the background rejection compared to the baseline by $\mathcal{O}$(10%). Any accumulated radiation damage also changes the cluster shape, reducing the signal efficiency compared to the baseline by $\sim$ 30 - 60%, but nearly all of the performance can be recovered through retraining of the network and updating the weights. Finally, the impact of noise was investigated, and retraining the network on noise-injected datasets was found to maintain performance within 6% of the baseline network trained and evaluated on noiseless data.

Shekar, Danush [Illinois U., Chicago]↗

PROTEUS: Machine Learning Driven Resilience for Extreme-scale Systems

The objective of this project is to design, develop, and evaluate scalable software to enhance resilience, data checkpointing, program restart, and analysis. The proposed tasks are to 1) develop scalable machine learning techniques to learn temporal change patterns in a scalable and in-situ manner, and to minimize data movement and maximize learning locally closest to data; 2) design a concise data representation and indexing mechanism to capture the distribution of changes in data that can guarantee point-wise user-defined tolerable errors while reducing the data storage requirements by an order of magnitude or more; 3) develop data reduction techniques as library modules; 4) exploit local SSD for minimizing data movement in storage hierarchy; 5) develop anomaly detection algorithms that can predict corruptions based on learning of emerging patterns; 6) develop software libraries to be incorporated within widely used data formats and APIs; and 7) evaluate the proposed software using DOE scientific applications. The outcomes of the proposed work are to satisfy many synergistic data reduction and resilience requirements for large-scale data intensive applications executed on extreme-scale computing systems. The developed mechanism for error-bound data approximation is directly applicable to existing scientific applications. Through machine learning from historical events and change distribution, this work will enable anomaly detection for DOE computer facility.

97 MATHEMATICS AND COMPUTING↗

200-IA-1 Operable Unit Human Health Risk and Kd Screen

The purpose of this environmental calculation file (ECF) is to provide the following: Document the data processing and data reduction steps taken to prepare the 200-IA-1 Operable Unit (OU) data set that will be used to calculate the sample-specific screening level human health risk evaluation; Document the data processing and data reduction steps taken to prepare the 200-IA-1 OU data set that will be used to identify analytes that could potentially impact groundwater in the future beneath the 200-IA-1 OU representative waste sites; Document the assumptions, equations, and methodologies used to calculate the screening levels for human health cancer risks and noncancer hazards for each representative waste site assigned to the 200-IA-1 OU. Individual measured soil concentrations from 0 to 4.6 m (15 ft) below ground surface (bgs) (shallow vadose zone) are used to calculate the total excess lifetime cancer risk (ELCR) and hazard index (HI) for the outdoor worker scenario to determine if there is a basis for remedial action. Individual measured soil concentrations from the ground surface to the groundwater table are used for the distribution coefficient (Kd) screen to identify analytes that could potentially impact groundwater in the future beneath the 200-IA-1 OU representative waste sites. This ECF supports DOE/RL-2020-51, 200-IA-1 OU Focused Feasibility Study, under the Comprehensive Environmental Response, Compensation, and Liability Act of 1980 (CERCLA). A risk characterization based upon the evaluation of the health risk estimates developed in this ECF will be presented in the focused feasibility study report.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

First Cosmology Results using Supernovae Ia from the Dark Energy Survey: Survey Overview, Performance, and Supernova Spectroscopy

We present details on the observing strategy, data-processing techniques, and spectroscopic targeting algorithms for the first three years of operation for the Dark Energy Survey Supernova Program (DES-SN). This five-year program using the Dark Energy Camera mounted on the 4 m Blanco telescope in Chile was designed to discover and follow supernovae (SNe) Ia over a wide redshift range (0.05 < z < 1.2) to measure the equation-of-state parameter of dark energy. We describe the SN program in full: strategy, observations, data reduction, spectroscopic follow-up observations, and classification. From three seasons of data, we have discovered 12,015 likely SNe, 308 of which have been spectroscopically confirmed, including 251 SNe Ia over a redshift range of 0.017 < z < 0.85. We determine the effective spectroscopic selection function for our sample and use it to investigate the redshift-dependent bias on the distance moduli of SNe Ia we have classified. The data presented here are used for the first cosmology analysis by DES-SN (“DES-SN3YR”), the results of which are given in Dark Energy Survey Collaboration et al. The 489 spectra that are used to define the DES-SN3YR sample are publicly available at https://des.ncsa.illinois.edu/releases/sn.

79 ASTRONOMY AND ASTROPHYSICS↗