Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scientific data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Elephants Sharing the Highway: Studying TCP Fairness in Large Transfers over High Throughput Links

Escalating bandwidth demand strains high-performance data networks, posing potential performance risks. TCP congestion control algorithms enhance reliability and optimize bandwidth usage. Network performance is influenced by factors such as AQM algorithms and router buffer size. In the context of constrained network resources, understanding how TCP flows share networks and the resulting performance impact is essential. This paper introduces insights into TCP fairness and performance involving a comparison of TCP CUBIC, Reno, Hamilton, and BBR versions 1 and 2 across real-world networks supporting high bandwidths of up to 25 Gbps. The research explores TCP behaviors with AQM algorithms like FIFO, FQ_CODEL, and RED, alongside diverse buffer sizes. Notably, findings reveal that manipulating buffers and queuing methods yields contrasting outcomes based on bandwidth. BBRv2 emerges as a superior fair algorithm, pivotal for swift transfers, particularly in scientific data scenarios. These results provide crucial guidance for future network design, ensuring equitable performance optimization.

Kiran, Mariam↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗

A comprehensive and fair comparison of two neural operators (with practical extensions) based on $\mathrm{FAIR}$ data

Neural operators can learn nonlinear mappings between function spaces and offer a new simulation paradigm for real-time prediction of complex dynamics for realistic diverse applications as well as for system identification in science and engineering. Herein, we investigate the performance of two neural operators, which have shown promising results so far, and we develop new practical extensions that will make them more accurate and robust and importantly more suitable for industrial-complexity applications. The first neural operator, DeepONet, was published in 2019 (Lu et al., 2019), and its original architecture was based on the universal approximation theorem of Chen & Chen (1995). The second one, named Fourier Neural Operator or FNO, was published in 2020, and it is based on parameterizing the integral kernel in the Fourier space. DeepONet is represented by a summation of products of neural networks (NNs), corresponding to the branch NN for the input function and the trunk NN for the output function; both NNs are general architectures, e.g., the branch NN can be replaced with a CNN or a ResNet. According to Kovachki et al. (2021), FNO in its continuous form can be viewed conceptually as a DeepONet with a specific architecture of the branch NN and a trunk NN represented by a trigonometric basis. In order to compare FNO with DeepONet computationally for realistic setups, we develop several extensions of FNO that can deal with complex geometric domains as well as mappings where the input and output function spaces are of different dimensions. We also develop an extended DeepONet with special features that provide inductive bias and accelerate training, and we present a faster implementation of DeepONet with cost comparable to the computational cost of FNO, which is based on the Fast Fourier Transform. Here we consider 16 different benchmarks to demonstrate the relative performance of the two neural operators, including instability wave analysis in hypersonic boundary layers, prediction of the vorticity field of a flapping airfoil, porous media simulations in complex-geometry domains, etc. We follow the guiding principles of FAIR (Findability, Accessibility, Interoperability, and Reusability) for scientific data management and stewardship. The performance of DeepONet and FNO is comparable for relatively simple settings, but for complex geometries the performance of FNO deteriorates greatly. We also compare theoretically the two neural operators and obtain similar error estimates for DeepONet and FNO under the same regularity assumptions.

42 ENGINEERING↗

Interdisciplinary Approaches Improve Understanding of Cryptogenic Species: A Historical Case Study of Crayfish in Montana, USA

ABSTRACT Cryptogenic species are those that are not yet “demonstrably native or introduced” in a given area, such as crayfish in western Montana, USA. Delving into evidence from diverse fields can help clarify the status of cryptogenic species. Primary historical sources and indigenous knowledge have informed various ecological questions but are seldom used to clarify the status of cryptogenic species, especially aquatic taxa. We summarize types of evidence used to illuminate species status and offer a case study applying historical sources to cryptogenic crayfish in Montana. We searched for crayfish mentions in dictionaries of indigenous languages and primary historical documents (e.g., travel journals, newspaper accounts, early biological surveys). Early explorer accounts we examined did not mention crayfish in Montana, though they noted them in the nearby Snake River drainage of Idaho and Wyoming. The first mention of crayfish in Montana that we found was from an 1868 newspaper account, probably referencing the upper Missouri River. The species was not noted, but an 1887 article suggested signal crayfish Pacifastacus leniusculus introductions near Bozeman, where non‐native signal crayfish still persist. The earliest crayfish mention we found west of the Great Divide in Montana was from 1934 in Ninepipes and probably referred to non‐native virile crayfish ( Faxonius virilis ). Signal crayfish were not mentioned west of the Divide until a 1944 newspaper account noted their 1937 introduction to the Bitterroot River drainage. The case study demonstrates the value of historical sources in identifying the absence of species and documenting their presence or introduction during periods predating scientific data.

Kennedy, Hampton L. [USDA Forest Service, Southern↗

The INTERSECT Open Federated Architecture for the Laboratory of the Future

A federated instrument-to-edge-to-center architecture is needed to autonomously collect, transfer, store, process, curate, and archive scientific data and reduce human-in-the-loop needs with (a) common interfaces to leverage community and custom software, (b) pluggability to permit adaptable solutions, reuse, and digital twins, and (c) an open standard to enable adoption by science facilities world-wide. The Selfdriven Experiments for Science/Interconnected Science Ecosystem (INTERSECT) Open Architecture enables science breakthroughs using intelligent networked systems, instruments and facilities with autonomous experiments, “self-driving” laboratories, smart manufacturing and artificial intelligence (AI) driven design, discovery and evaluation. It creates an open federated architecture for the laboratory of the future using a novel approach, consisting of (1) science use case design patterns, (2) a system of systems architecture, and (3) a microservice architecture.

Engelmann, Christian↗

Non-stationary precipitation design standards for stormwater infrastructure modernization at USAF installations

The resilience of defense infrastructure systems to a changing climate is critical for national security. Climate induced recurrent flooding is already impacting over 20 U.S. Air Force installations, underscoring the urgency of revisiting precipitation standards and stormwater infrastructure design. Despite growing scientific knowledge and an expanding set of tools for updating outdated precipitation standards based on the assumption of climate stationarity, the adoption of climate informed analyses remain limited in practice. This study utilizes an existing framework to update Intensity (or Depth)-Duration-Frequency (DDF) curves using an ensemble of future climate projections. Change factors in precipitation estimates are derived and applied to six USAF installations across the U.S. The analysis is further extended to evaluate the implications of climate-informed DDFs on stormwater infrastructure performance and flood analysis at Tyndall AFB. Results indicate that the current design precipitation estimates are likely to become obsolete in all six USAF bases by the end of the century. The wide range of change factors across 32 GCM ensembles highlights the need to integrate uncertainty and evolving scientific data into infrastructure planning. The study also finds that the impacts of a changing climate vary spatially and temporally, emphasizing the value of localized analysis for infrastructure decision-making. The work advances ongoing DoD and societal efforts to implement adaptation strategies aimed at enhancing infrastructure resilience.

Intensity-duration-frequency curves↗

Fourier-informed knot placement schemes for B-spline approximation

Fitting B-splines to scientific data is especially challenging when the given data contain noise, jumps, or corners. Here, we describe how periodic data sets with these features can be efficiently approximated with B-splines by analyzing the Fourier spectrum of the data. Our method uses a collection of spectral filters to compute high-order derivatives, smoothed versions of noisy data, and the locations of jump discontinuities. Further, these quantities are then combined to choose knots that capture the qualitative features of the data, leading to accurate B-spline approximations with few knots. The method we introduce is direct and does not require any intermediate B-spline fitting before choosing the final knot distribution. Aside from fast Fourier transforms to transfer to and from Fourier space, the method runs in linear time with very little communication. We assess performance on several test cases in one and two dimensions, including data sets with jump discontinuities and noise. These tests show the method fits discontinuous data without spurious oscillations and remains effective in the presence of noise.

97 MATHEMATICS AND COMPUTING↗

FAIR principles for AI models with a practical application for accelerated high energy diffraction microscopy

Abstract A concise and measurable set of FAIR (Findable, Accessible, Interoperable and Reusable) principles for scientific data is transforming the state-of-practice for data management and stewardship, supporting and enabling discovery and innovation. Learning from this initiative, and acknowledging the impact of artificial intelligence (AI) in the practice of science and engineering, we introduce a set of practical, concise, and measurable FAIR principles for AI models. We showcase how to create and share FAIR data and AI models within a unified computational framework combining the following elements: the Advanced Photon Source at Argonne National Laboratory, the Materials Data Facility, the Data and Learning Hub for Science, and funcX, and the Argonne Leadership Computing Facility (ALCF), in particular the ThetaGPU supercomputer and the SambaNova DataScale ® system at the ALCF AI Testbed. We describe how this domain-agnostic computational framework may be harnessed to enable autonomous AI-driven discovery.

47 OTHER INSTRUMENTATION↗

An ontology-based knowledge graph for representing interactions involving RNA molecules

The "RNA world" represents a novel frontier for the study of fundamental biological processes and human diseases and is paving the way for the development of new drugs tailored to each patient's biomolecular characteristics. Although scientific data about coding and non-coding RNA molecules are constantly produced and available from public repositories, they are scattered across different databases and a centralized, uniform, and semantically consistent representation of the "RNA world" is still lacking. We propose RNA-KG, a knowledge graph (KG) encompassing biological knowledge about RNAs gathered from more than 60 public databases, integrating functional relationships with genes, proteins, and chemicals and ontologically grounded biomedical concepts. To develop RNA-KG, we first identified, pre-processed, and characterized each data source; next, we built a meta-graph that provides an ontological description of the KG by representing all the bio-molecular entities and medical concepts of interest in this domain, as well as the types of interactions connecting them. Finally, we leveraged an instance-based semantically abstracted knowledge model to specify the ontological alignment according to which RNA-KG was generated. RNA-KG can be downloaded in different formats and also queried by a SPARQL endpoint. A thorough topological analysis of the resulting heterogeneous graph provides further insights into the characteristics of the "RNA world". RNA-KG can be both directly explored and visualized, and/or analyzed by applying computational methods to infer bio-medical knowledge from its heterogeneous nodes and edges. The resource can be easily updated with new experimental data, and specific views of the overall KG can be extracted according to the bio-medical problem to be studied.

59 BASIC BIOLOGICAL SCIENCES↗

Engage the Public in Science and Embrace Future Change with Human-Centric Stories, Art, and Imaginings

Human-centric stories that weave in real scientific data may be able to engage the public in environmental issues that do not yet directly affect them. Science fiction and artistically rendered futuristic scenarios can unleash the imagination and act as a lens to envision technological, social, and cultural aspects of transitioning to clean energy. Cultures with strong oral traditions use stories to record history, develop a shared identity, pass on environmental lessons, and prepare for future change. Through the lenses of our non-physical science disciplines (cultural anthropology and environmental sociology), we discuss challenges to communicating science, the use of narratives to illustrate ecosystem processes, and we report on a collaborative, interdisciplinary workshop and book project that is creating narratives of hope and visions for the future through inspiring art, short stories, and essays. We explore the surprising potential of 'cli-fi,' humor, games, and the research behind emotion and imagination-driven engagement. We describe methods that help people visualize their future well-being and we explore opportunities to spread those methods.

clean energy↗

Viricidal and bactericidal exciplex barrier-discharge lamps

A brief review is presented of investigations performed in 2002 – 2020 on ultraviolet inactivation of bacteria, vital cells, and viruses by using excilamps. The excilamp models that have been developed at the Institute of High-Current Electronics, SB RAS are briefly described. Scientific data acquired by now show that excilamps on KrCl{sup *}, KrBr{sup *}, and XeBr{sup *} molecules are an alternative to low-pressure mercury lamps with respect to optical parameters. These sources of optical radiation exhibit a bactericidal effect, and emission of KrCl and KrBr excilamps demonstrates viricidal action. The latter is actual due to expansion of coronavirus disease. (paper)

60 APPLIED LIFE SCIENCES↗

Numerical characterization of support recovery in sparse regression with correlated design

Sparse regression is employed in diverse scientific settings as a feature selection method. A pervasive aspect of scientific data is the presence of correlations between predictive features. These correlations hamper both feature selection and estimation and jeopardize conclusions drawn from estimated models. On the other hand, theoretical results on sparsity-inducing regularized regression have largely addressed conditions for selection consistency via asymptotics, and disregard the problem of model selection, whereby regularization parameters are chosen. In this numerical study, we address these issues through exhaustive characterization of the performance of several regression estimators, coupled with a range of model selection strategies. These estimators and selection criteria were examined across correlated regression problems with varying degrees of signal to noise, distributions of non-zero model coefficients, and model sparsity. Our results reveal a fundamental tradeoff between false positive and false negative control in all regression estimators and model selection criteria examined. Additionally, we numerically explore a transition point modulated by the signal-to-noise ratio and spectral properties of the design covariance matrix at which the selection accuracy of all considered algorithms degrades. Overall, we find that SCAD coupled with BIC or empirical Bayes model selection performs the best feature selection across the regression problems considered.

97 MATHEMATICS AND COMPUTING↗

Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models

Abstract We propose masked particle modeling (MPM) as a self-supervised method for learning generic, transferable, and reusable representations on unordered sets of inputs for use in high energy physics (HEP) scientific data. This work provides a novel scheme to perform masked modeling based pre-training to learn permutation invariant functions on sets. More generally, this work provides a step towards building large foundation models for HEP that can be generically pre-trained with self-supervised learning and later fine-tuned for a variety of down-stream tasks. In MPM, particles in a set are masked and the training objective is to recover their identity, as defined by a discretized token representation of a pre-trained vector quantized variational autoencoder. We study the efficacy of the method in samples of high energy jets at collider physics experiments, including studies on the impact of discretization, permutation invariance, and ordering. We also study the fine-tuning capability of the model, showing that it can be adapted to tasks such as supervised and weakly supervised jet classification, and that the model can transfer efficiently with small fine-tuning data sets to new classes and new data domains.

Heinrich, Lukas (ORCID:0000000240487584)↗

Running Ensemble Workflows at Extreme Scale: Lessons Learned and Path Forward

The ever-increasing volumes of scientific data combined with sophisticated techniques for extracting information from them have led to the increasing popularity of ensemble workflows which are a collection of runs of individual workflows. A traditional approach followed by scientists to run ensembles is to rely on simple scripts to execute different runs and manage resources. This approach is not scalable and is error-prone, thereby motivating the development of workflow management systems that specialize in executing ensembles on HPC clusters. However, when the size of both the ensemble and the target system reach extreme scales, existing workflow management systems face new challenges that hamper their efficient execution. In this paper, we describe our experience scaling an ensemble workflow from the computational biology domain from the early design stages to the execution at extreme scale on Summit, a leadership class supercomputer at the Oak Ridge National Laboratory. We discuss challenges that arise when scaling ensembles to several million runs on thousands of HPC nodes. We identify challenges with composition of the ensemble itself, its execution at large scale, post-processing of the generated data, and scalability of the file system. Based on the experience acquired, we develop a generic vision of the capabilities and abstractions to add to existing workflow management systems to enable the execution of ensemble workflows at extreme scales. We believe that the understanding of these fundamental challenges will help application teams along with workflow system developers with designing the next generation of infrastructure for composing and executing extreme-scale ensemble workflows.

Mehta, Kshitij↗

eCounter: Inline Per-IP Network Monitoring at Millisecond Resolution via eBPF

Scientific data acquisition (SciDAQ) systems are shifting from archive-based workflows to streaming paradigms, where real-time, fine-grained network monitoring becomes essential. While P4-enabled devices offer per-packet in-band observability, they require specialized switches and routers. Host-side tools like Prometheus exporters lack sufficient temporal granularity. To bridge this gap, we present eCounter, a lightweight, hardware-agnostic, inline telemetry agent built on extended Berkeley Packet Filter (eBPF). eCounter captures per-interface ingress and egress traffic, categorized by IP address and protocol, at millisecond to sub-millisecond resolution. In a 100 Gbps environment, it continuously exports up to 3,257 time-series bins per second with only 4% CPU utilization at a 35¿KiB/s data rate. We evaluate eCounter across diverse NIC MTU settings, hook types, CPU architectures and operating systems, and observed negligible impact on concurrent high-throughput streaming applications. Complexity analysis confirms that it can be readily scaled to distributed SciDAQ deployments.

Mei, Xinxin [Computational Sciences and Technology↗

ornladios/ADIOS2

ADIOS 2: The Adaptable Input Output (I/O) System version 2 is an open-source framework that addresses scientific data management challenges, e.g. scalable parallel I/O, as we approach the exascale era in high-performance computing (HPC). ADIOS 2 bindings are available in C++, C, Fortran, Python and can be used on supercomputers, personal computers, and cloud systems running on Linux, macOS and Windows. ADIOS 2 has out-of-the-box support for MPI and serial environments.

ECP↗

SurFE-XD (Surface curvature-driven Finite Elements model for Diffusion under eXtreme conditions)

SurFE-XD is a mesoscale finite element framework to model surface diffusion under mutliphysics environments. The code uses legacy C++ library dolphin wrapped with python in a FEniCS driven unified form language and just-in-time (JIT) compilation setting. The purpose of the release is to attract wide-ranging usage of the code along with publication supplementation to support reproducibility of scientific data. SurFE-XD has been originally conceived under the LDRD-DR funding for “High-Gradient (C-BAND) Breakdown tolerant accelerator materials project. Currently SrFE-XD support electrostatics and Thermo-elasticity driven surface diffusion kernels. Releasing the code will also enable to include contributions from other physical regimes e.g., plasticity and electrodynamics etc as well portability to GPU-based platforms.

Bagchi, Soumendu↗

CVEVOLVE

CVEvolve is an agentic AI system for autonomous algorithm discovery for scientific data processing. It creates workflows where large language model agents freely set up and configure development environments and evaluation harnesses, develop and improve data processing algorithms with designed exploration-exploitation balancing mechanisms, log history and findings in a structured database, and run holdout testing to ensure algorithm generalizability. CVEvolve offers a zero-code interface and does not require users to provide structured data and evaluation scripts.

Cherukara, MatthewJoseph [Argonne National Laborat↗