Engineering PapersSearch

SEARCH · Engineering Papers

Results for “BINARY DATA”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

X-ray and γ-ray beam interstellar communication and implications for SETI

The possibility of detecting artificial signals transmitted by alien civilizations via collimated X-ray or gamma-ray beams is investigated. The prospect of using such beams for human communication within the solar system and beyond is also discussed. Detector responses were simulated for input signals and analyzed using relative entropy. For simplicity, all signals were assumed to use on-off keying (OOK) modulation. “Real” signals were generated by taking digital files and sequentially feeding their raw binary data to the detector simulator, the resulting normalized information content of the detector signals was plotted and compared to random noise signals. Since jpeg files contain compressed information, these served as a proxy for artificial alien signals. This showed that there is a clear difference in measured information content between natural and artificial signals, even with relatively poor time resolution in the detector causing the signals to be smeared (dead-time/rise-time intervals many times longer than the duration between signal pulses). It was found that so long as the signal lasts for at least several rise-time/dead-time intervals, the distinction between random and artificial signals is obvious. A space-telescope with high time resolution for searching for such signals is briefly described and its basic requirements are outlined.

43 PARTICLE ACCELERATORS

Differential Seismic Phase Detection Probability as a Potential Discriminant of Explosions and Earthquakes

Deep learning models trained to estimate the probability of seismic P and S phases are rapidly expanding the scale of local event detections. Here, we evaluate the potential for deep learning model output phase detection probabilities to contribute to event‐type classification, particularly discrimination of single‐fired borehole explosions and earthquakes at local distances (<300 km). Motivated by the empirical success of P/S amplitude ratios, we consider the difference between P and S pick probability output from previously developed phase detection models, P prob −S prob ⁠, as a discriminant. Test data include M L ∼1–4 earthquakes and explosions observed by common seismographs in ten geologically diverse localities. Depending on the picking model and training data, binary classification using P prob −S prob with at least three stations can achieve approximately equivalent classification accuracy as P/S amplitude ratios without requiring any customization. Joint classification with P/S and P prob −S prob improves accuracy for most quality control scenarios. Pick probabilities are an efficient attribute to consider in explosion discrimination because they can be automated byproducts of event detection. They avoid the binary choice of picking or not picking weakly visible S waves common to explosions.

Duan, Chenglong [Rice Univ., Houston, TX (United S

Get Non-Real: Randomized Sketching for High-Dimensional Non-Real Valued Data (Final Report)

In our final report for DE-C0022186, we describe the work we did on this grant towards the goals we proposed. Our first goal was characterizing fundamental limits for sketching of discrete high-dimensional matrices with low-dimensional structures. Our second main goal was designing algorithms for data reconstruction from sketches. We focus on approaches that are either specifically designed for non-real-valued data (binary, finite field) or that will translate more readily to that setting.

97 MATHEMATICS AND COMPUTING

A New Vehicle-to-Vehicle Communication System: Visual-Enhanced Cooperative Traffic Operations

The advent of Connected and Autonomous Vehicles (CAVs) has highlighted the necessity for robust communication systems between vehicles and their environment. This study introduces a novel vehicle-to-vehicle (V2V) communication system, termed the Visual-Enhanced Cooperative Traffic Operations (VECTOR) system. The VECTOR system addresses the need for robust communication by converting dynamic data (including velocity and yaw angle data) into binary code, which is displayed on an LED panel mounted on the top of the vehicle. Following vehicles detect this panel and decode the information using a camera, implementing a visual-based communication method. VECTOR system employs a comprehensive five-module process. Initially, polynomial fitting techniques are applied to velocity data over fixed time intervals using third-degree polynomials, with validation via R² and MSE metrics. The second module converts velocity and yaw angle data into binary form, thereby enhancing detection and processing efficiency. The third module focuses on improving detection stability across various environmental conditions to enhance traffic safety. The fourth module decodes the binary data back into trajectory information, ensuring the fidelity of velocity and yaw angles. The final module integrates eco-control through the VECTOR system, employing advanced control algorithms to minimize energy consumption in CAVs. Experimental evaluations conducted using a modified CAV test platform based on the Lincoln MKZ demonstrate the feasibility and efficiency of the VECTOR system, achieving a 75% R-squared accuracy rate in replicating original velocity data. This methodology not only highlights potential applications but also underscores significant implications for advancing CAV technology.

Ma, Ke

Investigating the effects of local environment on nitrogen vacancies in high-entropy metal nitrides

High-entropy metal nitrides are an important material class in a variety of applications, and the role of nitrogen vacancies is of great importance for understanding their stability and mechanical properties. Here, we study six different high-entropy nitrides with eight different metal species to build a predictive model of the nitrogen-vacancy formation energy. We construct sets of supercells that maximize the number of unique nitrogen environments for a given chemistry, and then use density-functional theory to calculate the energy density for all nitrogen sites, and the vacancy formation energies for the highest, lowest, and a median subset based on the energy densities. The energy density of nitrogen sites correlates with the vacancy formation energies, for binary, ternary, and high-entropy nitrides. A linear regression model predicts the vacancy formation energies using only the nearest-neighbor composition; across our eight metals, we find the largest vacancy formation energies next to Hf, then Zr, Ti, V, Cr, Ta, Nb, and the lowest near Mo. Additionally, we see that binary nitride data show qualitatively similar vacancy formation energy trends for high-entropy nitrides; however, the binary data alone are insufficient to predict the complex nitride behavior. Our model is both predictive and easily interpretable, and correlates with experimental data.

DeSilva, Charith R. [Univ. of Illinois at Urbana-C

Blue Keanu: A Scientific Visualization Tool For Network Data

This software allows the user to visualize complex PCAP-ng files captured from network capture software such as Wireshark. The visualization runs in a GUI window that can be zoomed or moved to areas of interest in a waterfall type display. The user then can see an area of interest that looks different than the typical traffic visually, such as a human interaction or non-repetitive area of data. The program will tell the user the packet number and byte offset of interest for fast analysis of discrete atomic or non-random events. This is particularly useful for visualization of unknown binary format data, such as in PLC or SCADA protocols that may have human or other non-repetitive activity for further analysis, reverse engineering, or fast forensic analysis.

Durller, MichaelGeorge

From Existing and New Nuclear and Astrophysical Constraints to Stringent Limits on the Equation of State of Neutron-Rich Dense Matter

Through continuous progress in nuclear theory and experiment and an increasing number of neutron-star (NS) observations, a multitude of information about the equation of state (EOS) for matter at extreme densities is available. To constrain the EOS across its entire density range, this information needs to be combined consistently. However, the impact and model dependency of individual observations vary. Given their growing number, assessing the various methods is crucial to compare the respective effects on the EOS and discover potential biases. For this purpose, we present a broad compendium of different constraints and apply them individually to a large set of EOS candidates within a Bayesian framework. Specifically, we explore different ways of how chiral effective field theory and perturbative quantum chromodynamics can be used to place a likelihood on EOS candidates. We also investigate the impact of nuclear experimental constraints, as well as different radio and x-ray observations of NS masses and radii. This is augmented by reanalyses of the existing data from binary neutron star coalescences, in particular of GW170817, with improved models for the tidal waveform and kilonova light curves, which we also utilize to construct a tight upper limit of 2.39 M ⊙ on the TOV mass based on GW170817’s remnant. Our diverse set of constraints is eventually combined to obtain stringent limits on NS properties. We organize the combination in a way to distinguish between constraints where the systematic uncertainties are deemed small and those that rely on less conservative assumptions. For the former, we find the radius of the canonical 1.4 M ⊙ neutron star to be R 1.4 = 12.2 6 − 0.91 + 0.80 km and the TOV mass at M TOV = 2.2 5 − 0.22 + 0.42 M ⊙ (95% credibility). Including all the presented constraints yields R 1.4 = 12.2 0 − 0.48 + 0.50 km and M TOV = 2.3 0 − 0.20 + 0.07 M ⊙ . When comparing these limits to individual data points, we find that the quoted radius of HESS J1731-347 displays noticeable tension with other constraints. Constraining microphysical properties of the EOS proves more challenging. For instance, the symmetry energy slope is restricted to L sym = 48 − 25 + 21 MeV , where this constraint is mainly dominated by our reanalysis of the PREX-II and CREX experiment. Published by the American Physical Society 2025

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

FY24 Progress Report: SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking

Algorithms for Machine Learning (ML) and data analysis for the 3013 Surveillance Program have been developed in an ongoing collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). The objective of the algorithms is to automate the identification of corrosion and crack formation in the Inner Container Closure Weld Region (ICCWR) of the canister system used to store Pu-bearing material. Data for corrosion and cracking is collected from large binary files generated by a Laser Confocal Microscope (LCM), the Wide Area 3D Measurement System (WAMS), or,in a recent proposal, by a Scanning Electron Microscope (SEM). The ML software uses the physical attributes in the data files (e.g., one or more of: height, color, and 16-bit grayscale values as functions of position in a plane projection) to detect signs of surface corrosion and cracking after being trained on similar data, with the features to be detected. Although the initial scope included screening for broader indicators of corrosion, e.g., pitting, identification of potential cracks was prioritized for the past several years at the request of program leadership. Labeled training data is essential to developing the ML algorithm, and enhancements to data labeling capability have been developed to address this essential precursor to application of ML routines. Efficient labeling is particularly important in view of the large volume of data required to train ML algorithms and the relative rarity of cracks in the ICCWR data set. The updated program will read binary data from either LCM, WAMS or SEM files, interrogate data attributes, facilitate user labeling of data for training ML algorithms, execute ML algorithms, output parameters from trained ML algorithms, report ML model accuracy with respect to labeled data, and generate graphical representations for various analyses. In FY24, hourglass neural networks (HNNs) that were initiated in FY22 were further developed and tested using available LCM data, and their performance was tested against that of the alternative U-Net Neural Network algorithm structure. HNNs along with previously developed Convolutional Neural Networks (CNNs) and Deep Neural Networks (DNNs) comprise a suite of ML tools for identification of cracks in the ICCWR

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Transferable predictions of energetic and structural properties for refractory solid solution alloys across chemical compositions

We present a data-efficient approach to train graph neural networks (GNNs) on density functional theory (DFT) data for accurate and transferable predictions of energetic and structural properties of refractory solid solution alloys in the niobium-tantalum-vanadium (Nb-Ta-V) chemical space. We start by training the GNN model only on DFT data that describes refractory binary alloys niobium-tantalum (Nb-Ta), niobium-vanadium (Nb-V), and tantalum-vanadium (Ta-V) to predict formation enthalpy and root mean squared displacement. Once trained, the GNN predictions are tested on DFT data describing refractory ternary alloys Nb-Ta-V. While, unsurprisingly, direct transferability from binary to ternary is not sufficiently accurate, augmenting the training with only 1% of the available ternary data (uniformly distributed across the entire range of chemical compositions) improves significantly the quality of the GNN predictions. For comparison, we assess the transferability in the opposite direction by training GNN models on ternary Nb-Ta-V data and making predictions on binaries Nb-Ta, Nb-V, and Ta-V, which exhibits notably higher predictive errors. The proposed methodology, which favors transferability from lower-component to higher-component alloys, offers an efficient path towards avoiding the curse of dimensionality incurred when collecting DFT data for discovery and design of multi-component disordered alloys.

Density functional theory calculations

mzPeak: Designing a Scalable, Interoperable, and Future-Ready Mass Spectrometry Data Format

Advances in mass spectrometry (MS) instrumentation, such as higher resolution, faster scan speeds, and improved sensitivity, have significantly increased the volume and complexity of data. The growing adoption of imaging and ion mobility further amplifies these challenges across MS-based omics fields, including proteomics, metabolomics, and lipidomics. While these technologies unlock new possibilities, they also present significant challenges in data management, storage, and accessibility. Existing open formats, such as the XML-based community standards mzML and imzML, struggle to meet the demands of modern MS workflows due to their large file sizes, slow data access, and limited metadata support. Vendor-specific formats, while optimized for proprietary instruments, lack interoperability, comprehensive metadata support and long-term archival reliability. This white paper lays the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multi-dimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will enable enhanced interoperability across platforms, seamless support for complex workflows including ion mobility and MS imaging, and faster data access compared to existing community formats such as mzML. Its design will ensure data is managed in compliance with regulatory standards, essential for applications such as precision medicine and chemical safety, where long-term data integrity and accessibility are critical. For vendors, mzPeak provides a streamlined, open alternative to proprietary formats, reducing the burden of regulatory compliance while aligning with the industry's push for transparency and standardization. By offering a high-performance, interoperable solution, mzPeak positions vendors to meet customer demands for sustainable data management tools which will be able to handle emerging and future data types and workflows. mzPeak aspires to become the cornerstone of MS data management, empowering researchers, vendors, and developers to innovate and collaborate more effectively.

data formats

Barge Science Van IMU Data

This dataset contains high-frequency (10Hz) data from the GX5-45 IMU on the Barge Science vans. The data are all raw binary files.

17 WIND ENERGY

Raw Data

This dataset contains high-frequency (10Hz) data from the GX5-45 IMU on the Barge Science vans. The data are all raw binary files.

17 WIND ENERGY

HDBind: encoding of molecular structure with hyperdimensional binary representations

Traditional methods for identifying “hit” molecules from a large collection of potential drug-like candidates rely on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have a significant limitation in that they require exceptional computing capabilities for even relatively small collections of molecules. Increasingly large and complex state-of-the-art deep learning approaches have gained popularity with the promise to improve the productivity of drug design, notorious for its numerous failures. However, as deep learning models increase in their size and complexity, their acceleration at the hardware level becomes more challenging. Hyperdimensional Computing (HDC) has recently gained attention in the computer hardware community due to its algorithmic simplicity relative to deep learning approaches. The HDC learning paradigm, which represents data with high-dimension binary vectors, allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas (computer vision, bioinformatics, mass spectrometery, remote sensing, edge devices, etc.). To the best of our knowledge, our work is the first to consider HDC for the task of fast and efficient screening of modern drug-like compound libraries. We also propose the first HDC graph-based encoding methods for molecular data, demonstrating consistent and substantial improvement over previous work. We compare our approaches to alternative approaches on the well-studied MoleculeNet dataset and the recently proposed LIT-PCBA dataset derived from high quality PubChem assays. We demonstrate our methods on multiple target hardware platforms, including Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), showing at least an order of magnitude improvement in energy efficiency versus even our smallest neural network baseline model with a single hidden layer. Our work thus motivates further investigation into molecular representation learning to develop ultra-efficient pre-screening tools. We make our code publicly available at https://github.com/LLNL/hdbind.

59 BASIC BIOLOGICAL SCIENCES

Relativistic gas accretion onto supermassive black hole binaries from inspiral through merger

Accreting supermassive black hole binaries are powerful multimessenger sources emitting both gravitational and electromagnetic (EM) radiation. Understanding the accretion dynamics of these systems and predicting their distinctive EM signals is crucial to informing and guiding upcoming efforts aimed at detecting gravitational waves produced by these binaries. To this end, accurate numerical modeling is required to describe both the spacetime and the magnetized gas around the black holes. In this paper, we present two key advances in this field of research. First, we have developed a novel 3D general relativistic magnetohydrodynamics (GRMHD) framework that combines multiple numerical codes to simulate the inspiral and merger of supermassive black hole binaries starting from realistic initial data and running all the way through merger. Throughout the evolution, we adopt a simple but functional prescription to account for gas cooling through photon emission. Next, we have applied our new computational method to follow the time evolution of a circular, equal-mass, nonspinning black hole binary for ∼200 orbits, starting from a separation of 20⁢𝑟 𝑔 and reaching the postmerger evolutionary stage of the system. We have shown how mass continues to flow toward the binary even after the binary “decouples” from its surrounding disk, but the accretion rate onto the black holes diminishes. We have identified how the minidisks orbiting each black hole are slowly drained and eventually dissolve as the binary compresses. We confirm previous findings that the system’s luminosity decreases by a factor of a few during inspiral; however, we observe an abrupt increase by ∼50% in this quantity at the time of merger, likely accompanied by an equally abrupt change in spectrum. Lastly, we have demonstrated that during the inspiral, fluid ram pressure regulates the fraction of the magnetic flux transported to the binary that attaches to the black holes’ horizons.

Accretion disk & black-hole plasma

eDNAjoint: An R package for interpreting paired or semi‐paired environmental DNA and traditional survey data in a Bayesian framework

Abstract Environmental DNA (eDNA) sampling is increasingly used in surveys of species distribution as a potentially sensitive and efficient monitoring method. Yet access to modelling tools designed specifically for interpreting this new data type lags behind its ubiquity. While occupancy modelling software has dominated the analytical landscape for eDNA data analysis of single species, this type of model may not always be the most appropriate. The rate of eDNA detection often corresponds to species density, rather than just occupancy, and researchers often have access to observations from non‐genetic sampling methods at the same sites. To provide users access to a modelling framework designed to maximize the use of all available data, we developed an R package, eDNAjoint . The package provides an easy‐to‐use interface for fitting a ‘joint’ model that integrates data from paired or semi‐paired eDNA and traditional surveys in a Bayesian framework. The model can be used to estimate parameters like the probability of a false positive eDNA detection and mean catch rate at a site, and the package allows access to multiple model variations and Bayesian prior customization. Additional functionality can be used for model selection, summarising posteriors and comparing the relative sensitivities of the two survey methods. We demonstrate the use of eDNAjoint by fitting a variation of the model with site‐level covariates that scale the sensitivity of eDNA sampling relative to traditional sampling. The example workflow uses binary eDNA and seine count data for the endangered tidewater goby ( Eucyclogobius newberryi ) from a study by Schmelzle and Kinziger (2016). This use case includes a prior sensitivity analysis and an evaluation of the relationship between detection rates and environmental variables. eDNAjoint has the potential to greatly increase the range of users who will be able to rigorously analyse eDNA and traditional survey data in a Bayesian framework, understand if and how eDNA can improve monitoring practices, and gain confidence in the interpretability of eDNA data.

Keller, Abigail G. [Department of Environment Scie

Chandra Discovery of a Candidate Hyperluminous X-Ray Source in MCG+11-11-032

We present a multiwavelength analysis of MCG+11-11-032, a nearby active galactic nucleus (AGN), with a unique classification as being both a binary and a dual AGN candidate. With new Chandra observations, we aim to resolve any dual AGN system via imaging data and search for signs of a binary AGN via analysis of the X-ray spectrum. Analyzing the Chandra spectrum, we find no evidence of the previously suggested double-peaked Fe Kα lines; the spectrum is instead best fit by an absorbed power law with a single Fe Kα line, as well as an additional line centered at ≈7.5 keV. The Chandra observation reveals faint, soft, and extended X-ray emission, possibly linked to low-level nuclear outflows. Further analysis shows evidence for a compact hard source—MCG+11-11-032 X2—located 3.″3 from the primary AGN. Modeling MCG+11-11-032 X2 as a compact source, we find that it is relatively luminous (L2–10 keV=1.5−0.5+0.9×1041erg s$^{−1}$), and the location is coincident with a compact and off-nuclear source resolved in Hubble Space Telescope infrared (F105W) and optical (F621M, F547M) bands. Pairing our X-ray results with a 144 MHz radio detection at the host galaxy location, we observe X-ray and radio properties similar to those of ESO 243-49 HLX-1, suggesting that MCG+11-11-032 X2 may be a hyperluminous X-ray source. This detection with Chandra highlights the importance of a high-resolution X-ray imager as well as how previous binary AGN candidates detected with large-aperture instruments can benefit from high-resolution follow-up. Future spatially resolved optical spectra, and deeper X-ray observations, can better constrain the origin of MCG+11-11-032 X2.

79 ASTRONOMY AND ASTROPHYSICS

The influence of cloud cover on the reliability of satellite-based solar resource data

Satellite-based solar resource data are often developed and validated by using binary cloudiness categories: clear sky or overcast cloudy sky. To investigate the reliability of solar resource data in partially cloudy conditions, we estimate cloud fraction using two distinct algorithms: a physical retrieval model using surface observed global horizontal irradiance (GHI) and direct normal irradiance (DNI) and a temporal average of cloud mask data estimated by the observed DNI. Our analysis reveals a significant presence of scattered clouds, broken clouds, and mismatches between satellite- and surface-based cloud data at 17 surface sites across the contiguous United States, though confidently clear and cloudy conditions collectively account for more than 70 % of the data. Solar radiation is computed using the National Solar Radiation Database (NSRDB) algorithm and validated using surface observations. Here, our findings suggest that, in the presence of scattered clouds, NSRDB data for clear-sky conditions can be subject to significant overestimation. In cloudy-sky conditions classified by satellite data, DNI computed by the Fast All-sky Radiation Model for Solar applications with DNI (FARMS-DNI) can be underestimated when limited clouds are detected by surface observations. The bias observed in several cloudiness categories indicates that the NSRDB is exceptionally accurate in confidently clear conditions. However, clear-sky conditions with scattered clouds and mismatched cloud data contribute significantly to the overall uncertainties in the NSRDB. Therefore, future improvements in solar resource data should involve development and implementation of satellite-derived cloud fraction and should consider a novel radiative transfer model accounting for amplified cloud reflection. The evaluation within cloudiness categories also provides a physical rationale for the superior performance of FARMS-DNI compared to the Direct Insolation Simulation Code (DISC) in both cloudy-sky and all-sky conditions.

14 SOLAR ENERGY