Engineering PapersSearch

SEARCH · Engineering Papers

Results for “BINARY DATA”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

eDNAjoint: An R package for interpreting paired or semi‐paired environmental DNA and traditional survey data in a Bayesian framework

Abstract Environmental DNA (eDNA) sampling is increasingly used in surveys of species distribution as a potentially sensitive and efficient monitoring method. Yet access to modelling tools designed specifically for interpreting this new data type lags behind its ubiquity. While occupancy modelling software has dominated the analytical landscape for eDNA data analysis of single species, this type of model may not always be the most appropriate. The rate of eDNA detection often corresponds to species density, rather than just occupancy, and researchers often have access to observations from non‐genetic sampling methods at the same sites. To provide users access to a modelling framework designed to maximize the use of all available data, we developed an R package, eDNAjoint . The package provides an easy‐to‐use interface for fitting a ‘joint’ model that integrates data from paired or semi‐paired eDNA and traditional surveys in a Bayesian framework. The model can be used to estimate parameters like the probability of a false positive eDNA detection and mean catch rate at a site, and the package allows access to multiple model variations and Bayesian prior customization. Additional functionality can be used for model selection, summarising posteriors and comparing the relative sensitivities of the two survey methods. We demonstrate the use of eDNAjoint by fitting a variation of the model with site‐level covariates that scale the sensitivity of eDNA sampling relative to traditional sampling. The example workflow uses binary eDNA and seine count data for the endangered tidewater goby ( Eucyclogobius newberryi ) from a study by Schmelzle and Kinziger (2016). This use case includes a prior sensitivity analysis and an evaluation of the relationship between detection rates and environmental variables. eDNAjoint has the potential to greatly increase the range of users who will be able to rigorously analyse eDNA and traditional survey data in a Bayesian framework, understand if and how eDNA can improve monitoring practices, and gain confidence in the interpretability of eDNA data.

Keller, Abigail G. [Department of Environment Scie

Chandra Discovery of a Candidate Hyperluminous X-Ray Source in MCG+11-11-032

We present a multiwavelength analysis of MCG+11-11-032, a nearby active galactic nucleus (AGN), with a unique classification as being both a binary and a dual AGN candidate. With new Chandra observations, we aim to resolve any dual AGN system via imaging data and search for signs of a binary AGN via analysis of the X-ray spectrum. Analyzing the Chandra spectrum, we find no evidence of the previously suggested double-peaked Fe Kα lines; the spectrum is instead best fit by an absorbed power law with a single Fe Kα line, as well as an additional line centered at ≈7.5 keV. The Chandra observation reveals faint, soft, and extended X-ray emission, possibly linked to low-level nuclear outflows. Further analysis shows evidence for a compact hard source—MCG+11-11-032 X2—located 3.″3 from the primary AGN. Modeling MCG+11-11-032 X2 as a compact source, we find that it is relatively luminous (L2–10 keV=1.5−0.5+0.9×1041erg s$^{−1}$), and the location is coincident with a compact and off-nuclear source resolved in Hubble Space Telescope infrared (F105W) and optical (F621M, F547M) bands. Pairing our X-ray results with a 144 MHz radio detection at the host galaxy location, we observe X-ray and radio properties similar to those of ESO 243-49 HLX-1, suggesting that MCG+11-11-032 X2 may be a hyperluminous X-ray source. This detection with Chandra highlights the importance of a high-resolution X-ray imager as well as how previous binary AGN candidates detected with large-aperture instruments can benefit from high-resolution follow-up. Future spatially resolved optical spectra, and deeper X-ray observations, can better constrain the origin of MCG+11-11-032 X2.

79 ASTRONOMY AND ASTROPHYSICS

The influence of cloud cover on the reliability of satellite-based solar resource data

Satellite-based solar resource data are often developed and validated by using binary cloudiness categories: clear sky or overcast cloudy sky. To investigate the reliability of solar resource data in partially cloudy conditions, we estimate cloud fraction using two distinct algorithms: a physical retrieval model using surface observed global horizontal irradiance (GHI) and direct normal irradiance (DNI) and a temporal average of cloud mask data estimated by the observed DNI. Our analysis reveals a significant presence of scattered clouds, broken clouds, and mismatches between satellite- and surface-based cloud data at 17 surface sites across the contiguous United States, though confidently clear and cloudy conditions collectively account for more than 70 % of the data. Solar radiation is computed using the National Solar Radiation Database (NSRDB) algorithm and validated using surface observations. Here, our findings suggest that, in the presence of scattered clouds, NSRDB data for clear-sky conditions can be subject to significant overestimation. In cloudy-sky conditions classified by satellite data, DNI computed by the Fast All-sky Radiation Model for Solar applications with DNI (FARMS-DNI) can be underestimated when limited clouds are detected by surface observations. The bias observed in several cloudiness categories indicates that the NSRDB is exceptionally accurate in confidently clear conditions. However, clear-sky conditions with scattered clouds and mismatched cloud data contribute significantly to the overall uncertainties in the NSRDB. Therefore, future improvements in solar resource data should involve development and implementation of satellite-derived cloud fraction and should consider a novel radiative transfer model accounting for amplified cloud reflection. The evaluation within cloudiness categories also provides a physical rationale for the superior performance of FARMS-DNI compared to the Direct Insolation Simulation Code (DISC) in both cloudy-sky and all-sky conditions.

14 SOLAR ENERGY

Quantifying Epistemic Uncertainty in Binary Classification via Accuracy Gain

ABSTRACT Recently, a surge of interest has been given to quantifying epistemic uncertainty (EU), the reducible portion of uncertainty due to lack of data. We propose a novel EU estimator in the binary classification setting, as the posterior expected value of the empirical gain in accuracy between the current prediction and the optimal prediction. In order to validate the performance of our EU estimator, we introduce an experimental procedure where we take an existing dataset, remove a set of points, and compare the estimated EU with the observed change in accuracy. Through real and simulated data experiments, we demonstrate the effectiveness of our proposed EU estimator.

97 MATHEMATICS AND COMPUTING

System for controller area network payload decoding

A system for decoding an unknown automotive controller area network (“CAN”) message definitions. CAN data vehicle signal mappings are typically held in secret and varied by automotive model and year. Without knowledge of the mappings, the wealth of real-time vehicle data hidden in the automotive CAN packets is uninterpretable—impeding research, after-market tuning, efficiency and performance monitoring, fault diagnosis, and privacy-related technologies. This system can ascertain the CAN signals' boundaries (start bit and length), endianness (byte ordering), signedness (binary-to-integer encoding) from raw CAN data. This allows conversion of CAN data to time series. Interpreting the translated CAN data's physical meaning and finding a linear mapping to standard units (e.g., knowing the signal is speed and scaling values to represent units of miles per hour) can be achieved for many signals by leveraging diagnostic standards to obtain real-time measurements of in-vehicle systems. The system can be integrated into lightweight hardware enabling an OBD-II plugin for real-time in-vehicle CAN decoding or run on standard computers. The system can output a standard DBC file with the signal definition information.

Verma, Kiren E.

Vulcan-Forge: Architecture and Design of a Multi-Modal Forensic Analysis Plugin for CALDERA

Forge and VULCAN together describe an open-architecture cybersecurity analysis ecosystem that unifies forensic artifact processing, detection engineering, and vulnerability intelligence within integrated platforms. Forge operates as a plugin for MITRE CALDERA, ingesting diverse evidence formats—including EVTX, PCAP/PCAPNG, CSV, JSON, YAML, XML, binaries, and archives—to construct a unified artifact graph enriched with severity scoring, TLP classification, and audit trails. It provides subsystems for artifact parsing, streaming structured-data visualization, NetworkMiner-based packet inspection, PE/.NET binary analysis, and LLM-assisted triage and rule generation, with outputs validated against CCCS-YARA and pySigma schemas. VULCAN complements this by serving as a cybersecurity analyst platform that integrates a Neo4j knowledge graph, Qdrant vector retrieval, SSVC-based triage, and a local LLM to deliver CVE intelligence and forensic analysis through a multi-source ingest pipeline drawing from NVD, CISA KEV, EPSS, MITRE ATT&CK, and CAPEC. Together, they bridge structured threat intelligence with automated forensic analysis and detection workflows.

97 MATHEMATICS AND COMPUTING

NUM-DAT File Format Specification: Used in M-9 Gun Experiment Data Archiving

The M-9 Shock and Detonation Physics group executes experiments on gun and explosive platforms with large numbers of oscilloscopes used for data acquisition. The data acquisition from these oscilloscopes was automated many years ago using a custom piece of software called RunDig . The default save format from this software is a custom structure referred to as "NUM-DAT" format. This file format includes a text ".DAT" file which is a header file used to interpret the binary ".NUM" file which contains the oscilloscope data. The data save format was originally developed by John Vorthman and has been in use by M-9 personnel for over 20 years. This data format has been used for archiving data from experiments performed by M-9 personnel at TA-40, TA-39, and the TA-55 Impact Test Facility. Numerous custom analysis and visualization programs have also been developed, and continue to be used, that utilize this data format. This document describes the NUM-DAT format and provides code examples for reading the format and converting it to other formats.

47 OTHER INSTRUMENTATION

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES

SESAME: ASCII2 File Format

A new ASCII format for SESAME data is explicitly defined and dubbed ASCII2. This format fixes some of the onerous limitations of the legacy ASCII-styled format in addition to relaxing the FORTRAN style fixed-format layout of data tables. With the exception of the multi-tiered indexes of the legacy binary format, the new ASCII2 is more compatible with the binary in that it does not restrict the extent of various integer data (e.g., SESAME material and table numbers) to six digits nor does it restrict the extent or precision of the various floating point data. There will be very limited support of the legacy ASCII-styled formats in the future.

36 MATERIALS SCIENCE

Urban Parameters Arizona Urban Corridor 100m

132 Urban parameters based on building physical dimensions and location were generated for the cities in six Arizona Counties at 100m resolution using the NATURF model. To use the binary file with WRF, the binary file and the index file must be placed in their own directory in WRF_GEOG and accessed in the same way NUDAPT44 would be accessed.

Dumas, Melissa [ORNL] (ORCID:0000000233190846)

Data mining the missing ordered phases of Li/Na metal oxides

Data-driven discovery of Li-ion and Na-ion battery materials has been pioneered by generic materials data platforms such as the Materials Project. After decades of progress, it is timely to ask whether there remain underexplored compositional spaces. Here, in this work, we present a systematic data-mining effort to uncover missing ordered binary, ternary and quaternary Li/Na-containing metal oxides using high-throughput density functional theory (DFT). Building on 19,120 stable and metastable oxides entries from the Materials Project, we performed 13,245 additional calculations through isovalent substitutions of known ground states, experimentally reported compounds, and specific prototype structures. Our study identifies 36 new ground states within the GGA/GGA + U convex hull and 45 within the r 2 SCAN convex hull. Additionally, we identified 840 metastable compounds from GGA/GGA + U and 979 from r 2 SCAN that are absent in the present Materials Project databases. Moreover, we have tripled the metastable materials in compositional spaces with a molar ratio of cation/anion >1, highlighting the overlooked opportunities in this compositional space.

25 ENERGY STORAGE

Bistable random momentum transfer in a linear on-chip resonator

Optical switches and bifurcation rely on the nonlinear response of materials. Here, we demonstrate linear temporal bifurcation responses in a passive multimode microresonator, with strongly coupled chaotic and whispering gallery modes (WGMs). In microdisks, the chaotic modes exhibit broadband transfer within the deformed cavities, but their transient response is less explored and yields a random output of the analog signal distributed uniformly from “0” to “1.” Here, we build chaotic states by perturbing the multimode microring resonators with densely packed silicon nanocrystals on the waveguide surface. In vivo measurements reveal random and “digitized” output that ONLY populates around 0 and 1 intensity levels. The bus waveguide mode couples first to chaotic modes, then either dissipates or tunnels into stable WGMs. This binary pathway generates high-contrast, digitized outputs. In conclusion, the fully passive device enables real-time conversion of periodic clock signals into binary outputs with contrasts exceeding 12.3 dB, data rates of up to 10 7 · bits per second, and 20 dB dynamic range.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Validating sequential Monte Carlo for gravitational-wave inference

Nested sampling (NS) is the preferred stochastic sampling algorithm for gravitational-wave inference for compact binary coalescences. It can handle the complex nature of the gravitational-wave likelihood surface and provides an estimate of the Bayesian model evidence. However, there is another class of algorithms that meets the same requirements, but has not been used for gravitational-wave analyses: sequential Monte Carlo (SMC), an extension of importance sampling that maps samples from an initial density to a target density via a series of intermediate densities. In this work, we validate a type of SMC algorithm, called persistent sampling (PS), for gravitational-wave inference. We consider a range of different scenarios including binary black holes and binary neutron stars and real and simulated data and show that PS produces results that are consistent with NS whilst being, on average, 2 times more efficient and 2.74 times faster. This demonstrates that PS is a viable alternative to NS that should be considered for future gravitational-wave analyses.

black hole mergers

The Binary Fraction of Stars in the Dwarf Galaxy Ursa Minor via Dark Energy Spectroscopic Instrument

We utilize multi-epoch line-of-sight velocity measurements from the Milky Way Survey of the Dark Energy Spectroscopic Instrument to estimate the binary fraction for member stars in the dwarf spheroidal galaxy Ursa Minor. Our dataset comprises 670 distinct member stars, with a total of more than 2,000 observations collected over approximately one year. We constrain the binary fraction for UMi to be $0.61^{+0.16}_{-0.20}$ and $0.69^{+0.19}_{-0.17}$, with the binary orbital parameter distributions based on solar neighborhood observation from Duquennoy & Mayor (1991) and Moe & Di Stefano (2017), respectively. Furthermore, by dividing our data into two subsamples at the median metallicity, we identify that the binary fraction for the metal-rich ([Fe/H]>-2.14) population is slightly higher than that of the metal-poor ([Fe/H]<-2.14) population. Based on the Moe & Di Stefano model, the best-constrained binary fractions for metal-rich and metal-poor populations in UMi are $0.86^{+0.14}_{-0.24}$ and $0.48^{+0.26}_{-0.19}$, respectively. After a thorough examination, we find that this offset cannot be attributed to sample selection effects. We also divide our data into two subsamples according to their projected radius to the center of UMi, and find that the more centrally concentrated population in a denser environment has a lower binary fraction of $0.33^{+0.30}_{-0.20}$, compared with $1.00^{+0.00}_{-0.32}$ for the subsample in more outskirts.

Qiu, Tian [Shanghai Jiao Tong Univ. (China); DESI

FY25 Progress Report: SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking

Algorithms for Machine Learning (ML) and image analysis for the 3013 Surveillance Program have been developed in an ongoing collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). The objective of the algorithms is to automate the identification of corrosion and cracks in the Inner Container Closure Weld Region (ICCWR) of the canister system used to store Pu-bearing material. Data for corrosion and cracking is collected from large binary files generated by a Laser Confocal Microscope (LCM), the Wide Area 3D Measurement System (WAMS), or, in a recent proposal, by a Scanning Electron Microscope (SEM). The ML software uses the physical attributes in the data files (e.g., one or more of: height, color, and 16-bit grayscale values as functions of position in a plane projection) to detect signs of surface corrosion and cracking after being trained on similar data with the features to be detected. Although the initial scope included screening for broader indicators of corrosion, e.g., pitting, the identification of potential cracks was prioritized for the past several years at the request of program leadership.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Custom surface reflectance, shade mask, and equivalent water thickness maps for the Colorado Headwaters Ecological Spectroscopy Study (2025)

This dataset contains land surface reflectance estimates and additional derived products generated from NEON Imaging Spectrometer (NIS) data collected in the Upper Gunnison river basin during June and July of 2025. Data was collected over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). These products were derived from radiance and LiDAR data collected by the NEON Airborne Observation Platform (AOP) campaign funded by the Colorado Headwaters Ecological Spectroscopy Study (CHESS) (doi:10.15485/3017965). Products include per-pixel surface reflectance (rfl) and reflectance uncertainty (rfl_unc), observational data (obs), canopy equivalent water thickness (ewt), and shade masks. Atmospheric correction was performed per flightline using the ISOFIT (Imaging Spectrometer Optimal FITting) optimal estimation framework to estimate surface reflectance and the associated per-band reflectance uncertainty. Reflectance retrievals achieved a mean absolute error of 1.5% across diverse validation surfaces (see validation report.pdf). Equivalent water thickness was calculated from surface reflectance using the Beer–Lambert absorption of liquid water. Shade masks were generated based on the geometry between the sun angle, ground surface, and sensor at the time of flight. Data products are provided per-flightline and as mosaics for each domain. Flightline data products are provided as ENVI-formatted binary files (rfl, rfl_unc, ewt) and GeoTIFFs (shade). Reflectance and uncertainty mosaics are provided as tiled NetCDFs, while all other mosaicked products are provided as cloud-optimized GeoTIFFs. These formats are supported by common geospatial software (e.g., QGIS, ArcGIS, ENVI) and programmatic libraries in Python (e.g., rasterio, xarray, spectral, netCDF4) and R (e.g., terra, ncdf4). Processing workflows were designed to be equivalent to those used to generate the 2018 CHESS campaign airborne imaging spectroscopy data products (doi:10.15485/3013527). All outputs were co-registered to a common spatial grid to support time series analyses. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: Data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). Computational research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns

Thermodynamic modeling of aqueous acetic acid, butyric acid, and lactic acid solutions

Based on the activity coefficient – fugacity coefficient approach, a rigorous thermodynamic modeling study is presented for accurate correlation of vapor-liquid equilibrium data of aqueous solutions of acetic acid (293 to 391 K), butyric acid (325 to 436 K), lactic acid (378 to 409 K), and acetic acid + butyric acid binary mixture (358 to 421 K). In addition, the pH data of the three aqueous, single carboxylic acid solutions were measured at 298 to 328 K and successfully correlated. Given that these aqueous carboxylic acid solutions exhibit various degrees of association behavior in both vapor and liquid phases, the thermodynamic models considered for this study include the Redlich-Kwong equation of state (RK-EoS) and the Hayden-O’Connell equation of state (HOC-EoS) for the vapor phase fugacity coefficients and the electrolyte non-random two-liquid model (eNRTL) and the association electrolyte non-random two-liquid model (AeNRTL) for the liquid phase activity coefficients. The combination of the HOC-EoS for the vapor phase and the AeNRTL model for the liquid phase is found to provide the best correlation results, consistent with the fact that the HOC-EoS and the AeNRTL model explicitly account for association behaviors in the vapor phase and the liquid phase, respectively.

09 BIOMASS FUELS