Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “lossless”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Numerically exact configuration interaction at quadrillion-determinant scale

The combinatorial growth of configuration interaction (CI) has long limited this formally exact quantum chemistry method to only the smallest molecules. Here, we report a numerically exact CI calculation exceeding one quadrillion (10 15 ) determinants, made possible by a lossless categorical compression strategy within the small-tensor-product distributed active space (STP-DAS) framework. This approach overcomes the traditional memory bottlenecks of CI by a numerically exact compression of the wavefunction representation and reformulating the most computationally demanding matrix–vector operations. Using this method, we performed a fully relativistic CI calculation of the ground state of HBrTe with over 10 15 complex-valued determinants in just 34.5 h on 1000 computing nodes—the largest CI calculation ever reported. We further achieved fast computation for systems with hundreds of billions of determinants on only a few compute nodes. Extensive benchmarks confirm that the method retains full numerical exactness while cutting memory and computational cost by orders of magnitude. Compared to previous state-of-the-art CI calculations, this work achieves a 1000 times increase in CI space, a 10 6 -fold increase in floating-point operations performed, and a 10 6 -fold improvement in computational speed.

Computational chemistry↗

Ultimate precision limit of noise sensing and dark matter search

Abstract The nature of dark matter is unknown and calls for a systematical search. For axion dark matter, such a search relies on finding feeble random noise arising from the weak coupling between dark matter and microwave haloscopes. We model such process as a quantum channel and derive the fundamental precision limit of noise sensing. An entanglement-assisted strategy based on two-mode squeezed vacuum is thereby demonstrated optimal, while the optimality of a single-mode squeezed vacuum is found limited to the lossless case. We propose a “nulling” measurement (squeezing and photon counting) to achieve the optimal performances. In terms of the scan rate, even with 20-decibel of strength, single-mode squeezing still underperforms the vacuum limit which is achieved by photon counting on vacuum input; while the two-mode squeezed vacuum provides large and close-to-optimum advantage over the vacuum limit, thus more exotic quantum resources are no longer required. Our results highlight the necessity of entanglement assistance and microwave photon counting in dark matter search.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Real Time implementation of Artificial Intelligence compression algorithm for High-Speed Streaming Readout signals

The new generation of high-energy physics experiments plans to acquire data in streaming mode. With this approach, it is possible to access the information of the whole detector (organized in time slices) for optimal and lossless triggering of data acquisitions. With this approach, data rates, especially in large detectors, are often very high, and the network is likely to be the bottleneck for the entire Streaming Read Out system. The aim of this work is to study the implementation of a lossy compression algorithm based on Artificial Intelligence: an Autoencoder. With Machine Learning it is possible to achieve a high compression ratio and fast inference time with only a small degradation of the signals, almost negligible for the specific application. This work explores different configurations of the Autoencoder and the implementation on different hardware. Different Autoencoder configurations are explored to find the best trade-off between compression ratio and reconstruction loss, both for signals and energy spectrum. Different hardware implementations are also explored to find the best platform to achieve real-time performance for the specific application.

Rossi, Fabio (ORCID:0009000385713885)↗

A hybrid coupler for directing quantum light emission with high radiative Purcell enhancement to a dielectric metasurface lens

Quantum photonic technologies such as quantum sensing, metrology, and simulation could be transformatively enabled by the availability of integrated single photon sources with high radiative rates and photon collection efficiencies. We address these challenges for quantum emitters formed from color center defect sites such as those in hexagonal boron nitride, which are promising candidates as single photon sources due to their bright, stable, polarized, and room temperature emission. We report design of a nanophotonic coupler from color center quantum emitters to a dielectric metasurface lens. The coupler is comprised of a hybrid plasmonic–dielectric resonator that achieves a large radiative Purcell enhancement and partial control of far-field radiation. We report radiative Purcell factors up to 285 and photon collection efficiencies up to 89% for a lossless metasurface, applying a continuous hyperboloidal phase-front. Our hybrid plasmonic–dielectric coupler interfacing two nanophotonic elements is a compound optical element, analogous to those found in microscope objective lenses, which combine multiple optical functions into a single component for improved performance.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Creating large Fock states and massively squeezed states in optics using systems with nonlinear bound states in the continuum

The quantization of the electromagnetic field leads directly to the existence of quantum mechanical states, called Fock states, with an exact integer number of photons. Despite these fundamental states being long-understood, and despite their many potential applications, generating them is largely an open problem. For example, at optical frequencies, it is challenging to deterministically generate Fock states of order two and beyond. Here, we predict the existence of an effect in nonlinear optics, which enables the deterministic generation of large Fock states at arbitrary frequencies. The effect, which we call an n-photon bound state in the continuum, is one in which a photonic resonance (such as a cavity mode) becomes lossless when a precise number of photons n is inside the resonance. Based on analytical theory and numerical simulations, we show that these bound states enable a remarkable phenomenon in which a coherent state of light, when injected into a system supporting this bound state, can spontaneously evolve into a Fock state of a controllable photon number. This effect is also directly applicable for creating (highly) squeezed states of light, whose photon number fluctuations are (far) below the value expected from classical physics (i.e., shot noise). In conclusion, we suggest several examples of systems to experimentally realize the effects predicted here in nonlinear nanophotonic systems, showing examples of generating both optical Fock states with large n (n > 10), as well as more macroscopic photonic states with very large squeezing, with over 90% less noise (10 dB) than the classical value associated with shot noise.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Progress toward Accelogic compression in ROOT

For the last 7 years, Accelogic pioneered and perfected a radically new theory of numerical computing codenamed “Compressive Computing”, which has an extremely profound impact on real-world computer science [1]. At the core of this new theory is the discovery of one of its fundamental theorems which states that, under very general conditions, the vast majority (typically between 70% and 80%) of the bits used in modern large-scale numerical computations are absolutely irrelevant for the accuracy of the end result. This theory of Compressive Computing provides mechanisms able to identify (with high intelligence and surgical accuracy) the number of bits (i.e., the precision) that can be used to represent numbers without affecting the substance of the end results, as they are computed and vary in real time. The bottom-line outcome will be to provide state-of-the-art compression algorithms --and accompanying software libraries-- able to surpass the performance of the compression engines currently available in the ROOT [7] framework. The resulting technology has the capability to enable substantial economic and operational gains (including speedup) for High Energy and Nuclear Physics data storage/analysis. In our initial studies, a factor of nearly x4 (3.9) compression was achieved with RHIC/STAR data where ROOT compression managed only x1.4 [6].As a collaboration of experimental scientists, private industry, and the ROOT Team, our aim is to capitalize on the substantial success delivered by the initial effort and produce a robust technology properly packaged as an open-source tool that could be used by virtually every experiment around the world as means for improving data management and accessibility.In this contribution, we will present our efforts integrating our concepts of “functionally lossless compression” within the ROOT framework implementation, with the purpose of producing a basic solution readily integrated into HENP applications. We will also present our progress applying this compression through realistic examples of analysis from both the STAR and CMS experiments.

Canal, Ph.↗

A lightweight, user-configurable detector ASIC digital architecture with on-chip data compression for MHz X-ray coherent diffraction imaging

Today, most X-ray pixel detectors used at light sources transmit raw pixel data off the detector ASIC. With the availability of more advanced ASIC technology nodes for scientific application, more digital functionalities from the computing domains (e.g., compression) can be integrated directly into a detector ASIC to increase data velocity. In this paper, we describe a lightweight, user-configurable detector ASIC digital architecture with on-chip compression which can be implemented in 130 nm technologies in a reasonable area on the ASIC periphery. In addition, we present a design to efficiently handle the variable data from the stream of parallel compressors. The architecture includes user-selectable lossy and lossless compression blocks. The impact of lossy compression algorithms is evaluated on simulated and experimental X-ray ptychography datasets. This architecture is a practical approach to increase pixel detector frame rates towards the continuous 1 MHz regime for not only coherent imaging techniques such as ptychography, but also for other diffraction techniques at X-ray light sources.

47 OTHER INSTRUMENTATION↗

D-Egg: a dual PMT optical module for IceCube

The D-Egg, an acronym for "Dual optical sensors in an Ellipsoid Glass for Gen2," is one of the optical modules designed for future extensions of the IceCube experiment at the South Pole. The D-Egg has an elongated-sphere shape to maximize the photon-sensitive effective area while maintaining a narrow diameter to reduce the cost and the time needed for drilling of the deployment holes in the glacial ice for the optical modules at depths up to 2700 m. The D-Egg design is utilized for the IceCube Upgrade, the next stage of the IceCube project also known as IceCube-Gen2 Phase 1, where nearly half of the optical sensors to be deployed are D-Eggs. With two 8-inch high-quantum efficiency photomultiplier tubes (PMTs) per module, D-Eggs offer an increased effective area while retaining the successful design of the IceCube digital optical module (DOM). The convolution of the wavelength-dependent effective area and the Cherenkov emission spectrum provides an effective photodetection sensitivity that is 2.8 times larger than that of IceCube DOMs. The signal of each of the two PMTs is digitized using ultra-low-power 14-bit analog-to-digital converters with a sampling frequency of 240 MSPS, enabling a flexible event triggering, as well as seamless and lossless event recording of single-photon signals to multi-photons exceeding 200 photoelectrons within 10 ns. Mass production of D-Eggs has been completed, with 277 out of the 310 D-Eggs produced to be used in the IceCube Upgrade. In this paper, we report the design of the D-Eggs, as well as the sensitivity and the single to multi-photon detection performance of mass-produced D-Eggs measured in a laboratory using the built-in data acquisition system in each D-Egg optical sensor module.

47 OTHER INSTRUMENTATION↗

Singularities in nearly uniform one-dimensional condensates due to quantum diffusion

Dissipative systems often exhibit wavelength-dependent loss rates. One prominent example is Rydberg polaritons formed by electromagnetically induced transparency, which have long been a leading candidate for studying the physics of interacting photons and also hold promise as a platform for quantum information. In this system, dissipation is in the form of quantum diffusion, i.e., proportional to k 2 (k being the wavevector) and vanishing at long wavelengths as k→0. Here, we show that one-dimensional condensates subject to this type of loss are unstable to long-wavelength density fluctuations in an unusual manner: after a prolonged period in which the condensate appears to relax to a uniform state, local depleted regions quickly form and spread ballistically throughout the system. We connect this behavior to the leading-order equation for the nearly uniform condensate—a dispersive analog to the Kardar-Parisi-Zhang equation—which develops singularities in finite time. Furthermore, we show that the wavefronts of the depleted regions are described by purely dissipative solitons within a pair of hydrodynamic equations, with no counterpart in lossless condensates. We close by discussing conditions under which such singularities and the resulting solitons can be physically realized.

74 ATOMIC AND MOLECULAR PHYSICS↗

Ultrahigh supercurrent density in a two-dimensional topological material

Ongoing advances in superconductors continue to revolutionize technology thanks to the increasingly versatile and robust availability of lossless supercurrents. In particular, high supercurrent density can lead to more efficient and compact power transmission lines, high-field magnets, as well as high-performance nanoscale radiation detectors and superconducting spintronics. In this work, we report the discovery of an unprecedentedly high superconducting critical current density (17 MA / cm 2 at 0 T and 7 MA / cm 2 at 8 T) in 1T'–WS 2 , exceeding those of all reported two-dimensional superconductors to date. 1T'–WS 2 features a strongly anisotropic (both in- and out-of-plane) superconducting state that violates the Pauli paramagnetic limit signaling the presence of unconventional superconductivity. Spectroscopic imaging of the vortices further substantiates the anisotropic nature of the superconducting state. More intriguingly, the normal state of 1T'–WS 2 carries topological properties. The band structure obtained via angle-resolved photoemission spectroscopy and first-principles calculations points to a Z 2 topological invariant. The concomitance of topology and superconductivity in 1T'–WS 2 establishes it as a topological superconductor candidate, which is promising for the development of quantum computing technology.

36 MATERIALS SCIENCE↗

HPDR: High-Performance Portable Scientific Data Reduction Framework

The rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to 3.5× faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to 103 TB/s reduction throughput, providing up to 4× acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments.

Chen, Jieyang [University of Oregon]↗

Lossy Compression: An Online Multi-Stage Technology for High-Fidelity Synchro- Waveform Measurements

Effective real-time monitoring and analysis of distributed grids necessitate the use of synchro-waveform measurements, which capture almost all high-frequency disturbances and transient phenomena. However, due to limitations in high-speed measurements and network bandwidth, it is challenging to transfer all high-fidelity synchro-waveforms losslessly and successfully. To cope with these challenges, a hybrid-based online multi-stage compression algorithm is proposed to significantly improve the compression efficiency for synchro-waveform measurements. Initially, the multiple discrete Wavelet transformation is deployed to deconstruct the waveform components. The delta encoding is further developed to decrease the magnitude. In conjunction with the Lempel-Ziv-Markov chain, the hybrid compression algorithm is implemented to achieve real-time compression for the synchro-waveform measurements. Moreover, an innovative error index that synergizes the time and frequency domain error and correlation is formulated to evaluate the waveform distortion. By integrating compression ratio, suitable parameters can be optimally selected. Finally, the simulation, laboratory experiments, as well as field tests across a spectrum of sampling frequencies and time intervals are conducted to substantiate the efficacy of the proposed method. Here, the outcomes demonstrated that a compression ratio of approximately 15.5 and 17.83 can be reached for 0.5 s and 1 s data under both offline and online scenarios, which equates to a substantial 93.5% to 94.39% reduction in data storage requirements.

High-fidelity synchro-waveform measurements↗

OptZConfig: Efficient Parallel Optimization of Lossy Compression Configuration

Lossless compressors have very low compression ratios that do not meet the needs of today's large-scale scientific applications that produce vast volumes of data. Error-bounded lossy compression (EBLC) is considered a critical technique for the success of scientific research. Although EBLC allows users to set an error bound for the compression, users have been unable to specify the requirements on the compression quality, limiting practical use. Our contributions are: (1) We formulate the problem of configuring EBLC to preserve a user-defined metric as an optimization problem. This allows many classes of new metrics to be preserved, which improves over current practices. (2) We present a framework, OptZConfig, that can adapt to improvements in the search algorithm, compressor, and metrics with minimal changes, enabling future advancements in this area. (3) We demonstrate the advantages of our approach against the leading methods to configure compressors to preserve specific metrics. Here, our approach improves compression ratios against a specialized compressor by up to 3 x, has a 56x speedup over FRaZ, 1000x speedup over MGARD-QOI post tuning, and 110x speedup over systematic approaches which had not been bounded by compressors before.

97 MATHEMATICS AND COMPUTING↗

Origin of Soft-Switching Output Capacitance Loss in Cascode GaN HEMTs at High Frequencies

Output capacitance (C OSS ) loss (E DISS ) is produced when the C OSS of a power device is charged and discharged, which ideally should be a lossless process. This loss was recently revealed to be a crucial concern for GaN high electron mobility transistors (HEMTs) in high-frequency soft-switching applications. Among various GaN devices, the composite-type, cascode GaN HEMT was reported to show the largest E DISS with a voltage dependence distinct from discrete GaN HEMTs. However, the physical origins of the EDISS in cascode GaN HEMTs remain unclear. This work fills this gap by identifying three loss components and, for the first time, experimentally quantifying them in the multi-MHz resonant switching. These loss components include a) the avalanche loss of Si MOSFET, b) the intrinsic E DISS of GaN HEMT, and c) the Si avalanche-induced GaN turn-ON loss. The last component was found to dominate E DISS at high voltage. By eliminating the Si avalanche and the associated loss components (a) and (c), the E DISS of cascode GaN HEMTs can be reduced by up to 75% at the price of an increase in output charge and switching transition time. Furthermore, these results provide new physical insights and practical guidelines to trim the soft-switching loss of cascode GaN HEMTs in high-frequency applications.

42 ENGINEERING↗

Control Design of Passive Grid-Forming Inverters in Port-Hamiltonian Framework

This article presents a modified dispatchable virtual oscillator control approach for achieving the passivity of gridforming inverters (GFMs), without assuming constant voltage and constant frequency. The proposed control framework utilizes the Port-Hamiltonian (PH) based structure that mimics the behaviors of coupled harmonic oscillators, along with an energy ‘pumpingor-damping’ block and the Control by Interconnection (CbI) technique, to render the inverter passive. Once passivity is achieved, transient stability of the system will be guaranteed. The proposed control framework is composed of three loops: an outer power dispatching loop that generates the voltage and frequency references, a virtual oscillator loop that emulates the spontaneous synchronization of oscillators, and an inductor current loop that maintains lossless interconnection in PH systems. In conclusion, the study shows that the proposed control approach ensures the passivity of GFMs, facilitating the transient stability design of multi-inverter systems, as interconnections of passive systems remain passive and stable.

42 ENGINEERING↗

Enhancing ZFP: A Statistical Approach to Understanding and Reducing Error Bias in a Lossy Floating-Point Compression Algorithm

The amount of data generated and gathered in scientific simulations and data collection applications is continuously growing, putting mounting pressure on storage and bandwidth concerns. A means of reducing such issues is data compression; but, lossless data compression is typically ineffective when applied to floating-point data. Thus, users tend to apply a lossy data compressor, which allows for small deviations from the original data. It is essential to understand how the error from lossy compression impacts the accuracy of the data analytics. Thus, we must analyze not only the compression properties but the error as well. In this paper, we provide a statistical analysis of the error caused by ZFP compression, a state-of-the-art, lossy compression algorithm explicitly designed for floating-point data. We show that the error is indeed biased and propose simple modifications to the algorithm to neutralize the bias and further reduce the resulting error.

97 MATHEMATICS AND COMPUTING↗

HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs

Scientific applications produce vast amounts of data, posing grand challenges in the underlying data management and analytic tasks. Progressive compression is a promising way to address this problem, as it allows for on-demand data retrieval with significantly reduced data movement cost. However, most existing progressive methods are designed for CPUs, leaving a gap for them to unleash the power of today’s heterogeneous computing systems with GPUs.In this work, we propose HP-MDR, a high-performance and portable data refactoring and progressive retrieval framework for GPUs. Our contributions are four-fold: (1) We carefully optimize the bitplane encoding and lossless encoding, two key stages in progressive methods, to achieve high performance on GPUs; (2) We propose pipeline optimization and incorporate it with data refactoring and progressive retrieval workflows to further enhance the performance for large data process; (3) We leverage our framework to enable high-performance data retrieval with guaranteed error control for common Quantities of Interest; (4) We evaluate HP-MDR and compare it with state of the arts using five real-world datasets. Experimental results demonstrate that HP-MDR delivers an average 13.68 × and 6.31 × throughput in data refactoring and progressive retrieval tasks, respectively. It also leads to 11.22 × throughput for recomposing required data representations under Quantity-of-Interest error control and 6.04 × performance for the corresponding end-to-end data retrieval, when compared with state-of-the-art solutions.

Li, Yanliang [University of Oregon]↗

CODARcode/MGARD

MGARD is a software providing error-controlled lossy compression and data refactoring based on multi-grid theories. It transforms floating-point scientific data into a multilevel representation, followed by quantization and lossless encoding processes, resulting in a self-describing compressed buffer. It supports diverse data topologies, error control norms, and computing architectures.

Chen, Jieyang [University of Oregon]↗