Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data needs”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Proceedings for the Workshop on Applied Nuclear Data Activities 2025

The 2025 Workshop for Applied Nuclear Data Activities (WANDA) covered four topic areas in nuclear data: Nuclear Data and Deterrence, Nuclear Data Prioritization for Fusion, High-Assay Low-Enriched Uranium and Novel Moderators for Advanced Reactors, and Data Preservation and Data Workflows. The intention of this workshop is to connect different communities that are invested in nuclear data and have their own unique sets of needs for the purposes of sharing information, fostering collaboration in areas of shared interest, and leveraging synergistic capabilities. The attendance of federal program managers at these workshops is essential in creating awareness of the needs of their respective communities and in providing information to better guide funding investments. In each of these topical sessions, a general description of the nuclear data needs and/or capabilities was presented, along with discussions of existing capabilities that could be leveraged, potential synergistic needs or resources, and challenges that must be overcome for the application space to progress. There are many synergistic nuclear data needs among these application spaces. The discussion largely focused on increasing the accuracy of the nuclear data and better quantifying the data uncertainties that have the greatest impact on applications. The full-day session on data processing and data workflows was by nature intended to be synergistic and applicable to all technical sessions at WANDA. A notable common theme that was highlighted across all sessions was the need for accelerated delivery of nuclear data products across complex and time-consuming workflows.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

CMOS-Based Single-Cycle in-Memory XOR/XNOR

Big data applications are on the rise, and so is the number of data centers. The ever-increasing massive data pool needs to be periodically backed up in a secure environment. Moreover, a massive amount of securely backed-up data is required for training binary convolutional neural networks for image classification. XOR and XNOR operations are essential for large-scale data copy verification, encryption, and classification algorithms. The disproportionate speed of existing compute and memory units makes the von Neumann architecture inefficient to perform these Boolean operations. Compute-in-memory (CiM) has proved to be an optimum approach for such bulk computations. The existing CiM-based XOR/XNOR techniques either require multiple cycles for computing or add to the complexity of the fabrication process. Here, we propose a CMOS-based hardware topology for single-cycle in-memory XOR/XNOR operations. Our design provides at least 2× improvement in the latency compared with other existing CMOS-compatible solutions. We verify the proposed system through circuit/system-level simulations and evaluate its robustness using a 5000-point Monte Carlo variation analysis. This all-CMOS design paves the way for practical implementation of CiM XOR/XNOR at scaled technology nodes.

97 MATHEMATICS AND COMPUTING↗

CSAPR2 CMAC 2.0 level c1 data.

Raw data from ARM precipitation radars must be corrected for atmospheric phenomena and instrument characteristics (e.g., attenuation, clutter) to retrieve precipitation properties. The Corrected Moments in Antenna Coordinates Version 2 (CMAC2) value-added product (VAP) is a set of algorithms and code that does such corrections, and it also retrieves precipitation quantities from the radar measurements. Similar to the X-SAPR radars at the ARM main site, C-SAPR2 radar data also needs corrections and improvements. CMAC 2.0 has been updated to work with the ARM C-SAPR 2 radar. Data and fields that have been processed to: correct for velocity aliasing, unfold and generate a cross-polarimetric phase difference that is monotonically increasing, removing impulses caused by non-uniform beam filling and phase shift on backscatter, recalculate specific differential phase using a 20-point Sobel filter on the aforementioned phase, correct for liquid path attenuation using the polarimetric signals, estimate rainfall rates at the gate using the specific attenuation. In addition, CMAC writes the data out into a community-standard format netCDF File using the CF/Radial conventions. The data are therefore compatible with new and existing National Center for Atmospheric Research (NCAR) tools such as RadXConvert for converting to a variety of popular file formats.

54 ENVIRONMENTAL SCIENCES↗

Characterization and Analysis of the Energy-Reporting Accuracy of Connected Devices

Emerging energy-efficient building systems increasingly exhibit greater functionality, often requiring multiple operating modes (e.g. white-tunability for lighting products and data traffic for devices with networked, integrated sensors). This increased functionality makes energy consumption estimates more complex. Given that these functions consume energy, the energy performance of such building systems is dependent on what operating modes they use and how much time they spend in each mode. Devices and systems that can report their own energy consumption mitigate this energy-performance uncertainty. This study explores the energy-reporting accuracy of market-available connected electrical outlets. The study considers two residential-market products (five units each, one outlet per unit) and three commercial-market products (two units each, 18 to 24 outlets per unit) with the ability to report power drawn and/or energy consumed by devices connected to their receptacles. The products were purchased through typical market channels. Pacific Northwest National Laboratory (PNNL) conducted testing in December 2018 at its Connected Lighting Test Bed (CLTB), using a custom-developed test setup and method adapted from industry standards. The setup collected energy-consumption data reported by the outlet devices under test (DUTs) at one-minute intervals and compared that data with measurements taken by a reference meter over a range of test conditions. The residential products reported power draw but not interval or cumulative energy consumption. The commercial products reported both power draw and cumulative energy consumption. Relative reporting error (RRE) was calculated for all measurements, and analysis of the results revealed variations across devices and test conditions. The total number of measurements (50 for each residential product, 60 for each commercial product) offers an appreciable comparison of performance at the make/model level. The average RRE of the residential products derived from reported power draw was -0.02% and -1.20%. The average RRE for two of the three the commercial products derived from reported power draw was worse than those of the residential products (-2.40%, -2.72%, -0.36%). The internal integration of power over time, used to calculate cumulative energy consumption, typically occurs at current and voltage sampling rates much higher than once per minute. This suggests that the average commercial-product RRE derived from reported energy consumption should be very consistent and better than performance based on reported power draw. However, the RRE derived from reported energy consumption varied significantly across the three makes of commercial-market products and was uniformly less accurate than performance based on reported power draw. Subsequent analysis identified a number of root causes for this decrease in performance, most of which were related to reporting resolution. The goals of this study are to generate awareness of building systems capable of reporting their own energy consumption, further interest in the value of energy data for a variety of uses, draw attention to how the accuracy of reported metrics can be characterized, and quantify the performance variation found in marketavailable products. The results of this study and subsequent related work may be relevant to stakeholders in industry-specification and standards-development organizations. The methods this study employs could inform test and measurement procedures and performance classifications for connected outlets, lighting products, and other building systems capable of reporting their own energy consumption. The study concludes with stakeholder recommendations, including the following: • Energy-reporting device and system manufacturers developing products that report energy consumption should characterize the accuracy of reported metrics using a reference meter calibrated by an independent laboratory that was accredited by an ILAC MRA signatory (and whose scope of accreditation explicitly covers energy measurement), and should include this information on product data sheets. • Standards and specification development organizations should develop application-specific performance classifications that end users can understand and relate to their energy-data use needs (e.g., 2% accuracy class for utility streetlight energy billing needs, or 10% accuracy class for ESCO performance verification needs). • Current or potential owners, operators, and specifiers of energy-reporting building systems should rigorously analyze the dependency of current and planned energy-data use cases on accuracy, noting in particular the dependence (or lack thereof) on relative vs. absolute accuracy, and on trueness vs. precision (i.e., repeatability), and should communicate use-case needs to industry standards and specification organizations.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Characterization and Analysis of the Energy-Reporting Accuracy of Connected Devices

Emerging energy-efficient building systems increasingly exhibit greater functionality, often requiring multiple operating modes (e.g. white-tunability for lighting products and data traffic for devices with networked, integrated sensors). This increased functionality makes energy consumption estimates more complex. Given that these functions consume energy, the energy performance of such building systems is dependent on what operating modes they use and how much time they spend in each mode. Devices and systems that can report their own energy consumption mitigate this energy-performance uncertainty. This study explores the energy-reporting accuracy of market-available connected electrical outlets. The study considers two residential-market products (five units each, one outlet per unit) and three commercial-market products (two units each, 18 to 24 outlets per unit) with the ability to report power drawn and/or energy consumed by devices connected to their receptacles. The products were purchased through typical market channels. Pacific Northwest National Laboratory (PNNL) conducted testing in December 2018 at its Connected Lighting Test Bed (CLTB), using a custom-developed test setup and method adapted from industry standards. The setup collected energy-consumption data reported by the outlet devices under test (DUTs) at one-minute intervals and compared that data with measurements taken by a reference meter over a range of test conditions. The residential products reported power draw but not interval or cumulative energy consumption. The commercial products reported both power draw and cumulative energy consumption. Relative reporting error (RRE) was calculated for all measurements, and analysis of the results revealed variations across devices and test conditions. The total number of measurements (50 for each residential product, 60 for each commercial product) offers an appreciable comparison of performance at the make/model level. The average RRE of the residential products derived from reported power draw was -0.02% and -1.20%. The average RRE for two of the three the commercial products derived from reported power draw was worse than those of the residential products (-2.40%, -2.72%, -0.36%). The internal integration of power over time, used to calculate cumulative energy consumption, typically occurs at current and voltage sampling rates much higher than once per minute. This suggests that the average commercial-product RRE derived from reported energy consumption should be very consistent and better than performance based on reported power draw. However, the RRE derived from reported energy consumption varied significantly across the three makes of commercial-market products and was uniformly less accurate than performance based on reported power draw. Subsequent analysis identified a number of root causes for this decrease in performance, most of which were related to reporting resolution. The goals of this study are to generate awareness of building systems capable of reporting their own energy consumption, further interest in the value of energy data for a variety of uses, draw attention to how the accuracy of reported metrics can be characterized, and quantify the performance variation found in marketavailable products. The results of this study and subsequent related work may be relevant to stakeholders in industry-specification and standards-development organizations. The methods this study employs could inform test and measurement procedures and performance classifications for connected outlets, lighting products, and other building systems capable of reporting their own energy consumption. The study concludes with stakeholder recommendations, including the following: • Energy-reporting device and system manufacturers developing products that report energy consumption should characterize the accuracy of reported metrics using a reference meter calibrated by an independent laboratory that was accredited by an ILAC MRA signatory (and whose scope of accreditation explicitly covers energy measurement), and should include this information on product data sheets. • Standards and specification development organizations should develop application-specific performance classifications that end users can understand and relate to their energy-data use needs (e.g., 2% accuracy class for utility streetlight energy billing needs, or 10% accuracy class for ESCO performance verification needs). • Current or potential owners, operators, and specifiers of energy-reporting building systems should rigorously analyze the dependency of current and planned energy-data use cases on accuracy, noting in particular the dependence (or lack thereof) on relative vs. absolute accuracy, and on trueness vs. precision (i.e., repeatability), and should communicate use-case needs to industry standards and specification organizations.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2021 NIRT Mini Drill (After Action Report)

During the summer and fall of 2021, several functional area drills were held that focused on exercising Consequence Management’s (CM) ability to extract and use data from RadResponder for the purpose of answering intermediate-phase questions presented as technical inject requests for information (RFI) in Sandia National Laboratories (SNL) Consequence Management Operational System (COSMOS) software. The scenario chosen was that of Northern Lights 2016 (NL16) which was a large-scale nuclear power plant (NPP) release exercise in the state of Minnesota. The NL16 data was extracted from the Radiological Assessment and Monitoring System (RAMS) event where it was created and was reformatted for implanting to a new RadResponder event. Next, the beta-version of a laboratory sample data simulator was used to generate more sample data that was injected to the event. Five “mini-drills” were devised with each prompt defined by a data-based need. For each drill, a team of assessment and NARAC scientists worked the problem using the drill prompt and the available data in RadResponder. The teams held a kickoff meeting, had several days to work the problem, and then reported their results as well as observations in a hotwash. Several areas for improvement in both the software and process were identified during the course of these drills. This report will document the process of addressing each RFI and the discovered gaps in both software capability and methodology so that they can be considered for future development and investment by the CM and NIRT programs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

CritView User’s Guide Rev. 2

This document serves as a user’s guide for the CritView code, version 1.05. It supersedes the previous revision (Rev. 1), which was applicable to version 1.04 of CritView. This release of CritView also includes version 1.09 of the database, which replaces version 1.08. The CritView code is used as an electronic equivalent of a nuclear criticality handbook (e.g., ARH-600). This code takes an electronic data library and allows the user to plot data as needed. This approach has two distinct advantages over a paper handbook. First, the database can be easily expanded to include additional data sources (e.g., other handbooks, configurations, or modeling techniques). Secondly, the code provides flexibility by allowing the user to easily change the units and parameters of the plots.

Finfrock, Scott H. [Savannah River Nuclear Solutio↗

EXFOR-NSR PDF database: a system for nuclear knowledge preservation and data curation

Current needs of nuclear science and technology include complete, well-documented, and easily verifiable nuclear data. The complete data records require supporting nuclear bibliography, presently stored in dedicated libraries, in addition, to actual data. Additionally, experimental nuclear reaction data (EXFOR) and Nuclear Science References (NSR) databases contain compilations based on primary (journals) and secondary (conference proceedings, theses, preprints, etc.) publications, and data received from authors via private communications. The secondary library materials and private communications often represent a bottleneck for nuclear data verification, compilation, evaluation, and dissemination activities. To address this issue, bibliographic materials were scanned into PDF (Portable Document Format) files and uploaded in a relational database. The traditional scope of nuclear databases that includes meta-data and numbers derived from data in specialized formats was broadened to accommodate the large volumes of original nuclear data publications. The complete PDF publication files were stored in a relational database as Binary Large OBjects (BLOB). This unique collection of nuclear data compilations and supporting publications generate many opportunities for machine learning applications. The Web interfaces for authorized and public access to the EXFOR-NSR nuclear publications database were implemented at the U.S. National Nuclear Data Center, https://www.nndc.bnl.gov/ and IAEA Nuclear Data Section, https://www-nds.iaea.org/ . The current system is complementary to major nuclear libraries and narrowly focused on nuclear data compilation and evaluation procedures. The contents of the PDF database, details of implementation, and Web interface are described. New capabilities for data curation, knowledge preservation, worldwide dissemination, and natural language processing (NLP) applications are given.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Power System Waveform Datasets for Machine Learning

The desire for increased visibility across the electricity grid will necessarily increase the deployment of sensing and measurement devices and associated data management needs to unprecedented levels. For the existing sensing and measurement infrastructure, there remains a great amount of “value” yet to be extracted through advanced data management and analytics. Availability of more data will not, by itself, lead to changes in grid visibility, security, and resiliency. To create the predictive and prescriptive environment required to enable new markets and transactions for customer revenue and a reliable grid, the data must be collected, organized, evaluated, and analyzed using sophisticated algorithms to provide actionable information allowing operators and customers to reliably manage an increasingly complex grid. Progress in artificial intelligence (AI) has been largely driven by large, publicly available datasets that can be used to train AI algorithms such as MNIST, a database of handwritten images of digits, and ImageNet, an image database of everyday objects. These types of publicly available databases of real-world training datasets have been largely credited for advancement of image processing, computer vision, and deep learning algorithms that these use cases deploy. However, in the power systems industry to date, there are few databases with proper event labeling, and data access to a publicly available collection of power system event waveforms that will allow users to interact with grid signature data. Publicly available datasets of power system event waveforms, such as the DOE/EPRI dataset, often lack critical metadata or contain limited examples of each event type, and data formats vary widely across these datasets.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Adaptation of Temperature Dependent Thermal Neutron Libraries for Fast On-The-Fly Monte Carlo Sampling at Arbitrary Temperatures

Thermal neutron scattering data is necessary for reactor simulations of thermal reactors, but the data storage required can be unwieldy. This is exacerbated when simulations require a large temperature range, as is the case for accident scenarios and reactor start-up. Previous work conducted at RPI has explored and demonstrated a proof of concept of On-The-Fly (OTF) data storage with light water and crystalline graphite based on ENDF/B-VII.1. This provides an updated data storage paradigm for incoherent inelastic cross-sections that allow the data to be used to generate necessary sampling Cumulative Density Functions (CDFs) for any desired temperature. The work presented in this paper extends the previous work by implementing the proof of concept based on ENDF/B-VIII.0, instead of ENDF/B-VII.1, and to include additional materials, such as beryllium and oxygen in beryllium oxide and the new reactor grades of porous graphite. Completion of this work allows a single thermal data file to represent any thermal scattering event for a given material OTF for the temperature of interest, which greatly reduces the data storage needed with only a marginal cost in the calculation of the CDFs OTF. This paper will review the basic methodologies in generating the OTF data, as well as present the updates and improvements made in the generation process. Application of the methodologies covered in this paper show a significant reduction in the storage requirements for inelastic incoherent scattering events. In which, the evaluated ENDF/B-VIII.0 thermal scattering library can be reduced from 563.1 MB to 10.9 MB for light water and 339.3 MB to 13.6 MB for Be in BeO.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue↗

CEAZ: Accelerating Parallel I/O Via Hardware-Algorithm Co-Designed Adaptive Lossy Compression

As supercomputers continue to grow to exa-scale, the amount of data that needs to be saved or transmitted is exploding. To this end, many previous works have studied using error-bounded lossy compressors to reduce the data size and improve the I/O performance. However, little work has been done for effectively offloading lossy compression onto FPGA-based SmartNICs to reduce the compression overhead. In this paper, we propose a hardware-algorithm co-design of efficient and adaptive lossy compressor for scientific data on FPGAs (called CEAZ) to accelerate parallel I/O. Our contribution is fourfold: (1) We propose an efficient Huffman coding approach that can adaptively update Huffman codewords online based on codewords generated offline (from a variety of representative scientific datasets). (2) We derive a theoretical analysis to support a precise control of compression ratio under an error-bounded compression mode, enabling accurate offline Huffman codewords generation. This also help us create a fixed-ratio compression mode for consistent throughput. (3) We develop an efficient compression pipeline by adopting cuSZ’s dual-quantization algorithm to our hardware use case. (4) We evaluate CEAC on five real-world datasets with both a single FPGA board and 256 nodes from Bridges2 supercomputer. Experiments show that CEAZ outperforms the second-best FPGA-based lossy compressor by 2× of throughput and 9.6× of compression ratio. It also improves MPI_File_write and MPI_Gather throughputs by up to 32.7× and 31.4×, respectively.

Zhang, Chengming↗

Image processing workflow yielding high contrast synchrotron nanoscale computed tomography data from Ni-YSZ electrodes

The operating lifetime of Ni-YSZ fuel electrodes used in solid oxide electrolysis cells and fuel cells (SOECs and SOFCs) is limited by Ni redistribution, one of the primary degradation mechanisms that must be overcome to extend the longevity and maximize the performance of SOECs and SOFCs. To achieve this, 3D microstructural data is needed to relate both initial performance and performance loss over time to microstructural properties and their evolution throughout operation under various conditions. However, 3D microstructure data remains relatively scarce within the literature due to multiple challenges in acquiring and analyzing such data reliably. This work presents a workflow for acquiring and processing synchrotron X-ray nanoscale computed tomography (nano-CT) data from Ni-YSZ electrodes. Parameters for each step in the nano-CT workflow are described up to the final result (a 3D reconstruction), with particular emphasis on image alignment using freely available software. Following the results of a parametric sweep of the image alignment step, high contrast, low signal-to-noise 3D nano-CT data is obtained with relatively short compute times. While the exact methods best suited to samples with different microstructural qualities, or similar Ni-YSZ nano-CT data obtained from other sources may deviate from the solution found herein, this work also generalizes the decision points and evaluation of each step to provide a starting point to adapt this workflow to other datasets.

08 HYDROGEN↗

Approach to using 3D laser-induced breakdown spectroscopy (LIBS) data to explore the interaction of FLiNaK and FLiBe molten salts with nuclear-grade graphite

Nuclear graphite has historically been a key component of many nuclear reactor designs and has emerged as key to numerous advanced nuclear reactor design concepts. Molten salt reactors (MSRs) are one broad group of advanced reactor designs currently being pursued by industry for commercialization. Several MSR designs under consideration use graphitic materials that directly interface with a molten salt, whether it is a fuel salt, coolant salt, or both. Therefore, the interaction of graphite materials with molten salts must be understood. To gain this required understanding, a range of data is needed including porosity, strength, and composition as a function of different salt exposure parameters. In this study, a laser-induced breakdown spectroscopy (LIBS) measurement and data analysis methodology was developed to obtain spatially resolved elemental composition information for graphite samples exposed to a molten fluoride salt. Traditional univariate emission line analysis of atomic, ionic, and molecular optical emission signals was coupled via correlation analysis with spectral decomposition of the data using principal component analysis. Elemental depth profiling and elemental mapping were also performed to visualize salt–graphite interactions. LIBS was demonstrated to be useful for measuring key analytes such as fluorine and hydrogen, which are troublesome for other analysis techniques. Evidence for complex behavior was found, thereby demonstrating the usefulness of the developed approach for future systematic studies.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Opportunities for an Integrated Web-Based Workbench for Data Access and Analysis - 20244

Management of environmental issues can require integration of multiple types of data and information, conducting data analysis and interpretation, and providing data visualization for effective communications. These data elements are important for site management to support regulator interactions and provide defensibility for remedial decisions. Databases and information repositories are core elements of managing data; however, efficient data access and analysis also enable effective site management. The U.S. Department of Energy (DOE) Hanford Site is an example of a complex site with a voluminous quantity of environmental data and a need for efficient site management. Different tiers of data and information tools have been developed and deployed to address site needs. These tools are configured for ready access via the web site interfaces and meet the rigorous quality requirements for environmental site management. Evolving efforts are focused on an integrated platform to meet site environmental management needs. In this platform, users can access site information at multiple levels of detail based on their need and permissions, so that data and associated analyses are presented within the context of the site mission and the user's management or technical needs. This concept is not only applicable at individual sites like Hanford but also applicable at other sites within the DOE complex. An integrated web-based architecture that links data visualization, data analytics, and management tools can provide holistic access to large data sets, minimize complexity, and maximize interactivity and technical communication. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Recommendations for Radiological Data Assessment Implementation

Since thorough verification, validation, and data quality assessment (DQA) processes may delay incident commanders and elected and appointed officials from making key decisions, the U.S. Department of Homeland Security (DHS) National Urban Security Technology Laboratory (NUSTL) tasked Pacific Northwest National Laboratory (PNNL) to develop tools and guidance for the federal, state, local, tribal, and territorial (FSLTT) responders using research, discussions with select FSLTT responders, and the experiences of subject matter experts in incident response. This report provides NUSTL with recommendations on best practices for verification, validation, and data quality assessment for the data collected by responders during a radiological or nuclear event. The purpose of this report is to inform the development of a DQA toolkit aimed at the needs of data assessors during the response to a radiological incident.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Streaming readout for next generation electron scattering experiments

Current and future experiments at the high-intensity frontier are expected to produce an enormous amount of data that needs to be collected and stored for offline analysis. Thanks to the continuous progress in computing and networking technology, it is now possible to replace the standard ‘triggered’ data acquisition systems with a new, simplified and outperforming scheme. ‘Streaming readout’ (SRO) DAQ aims to replace the hardware-based trigger with a much more powerful and flexible software-based one, that considers the whole detector information for efficient real-time data tagging and selection. Considering the crucial role of DAQ in an experiment, validation with on-field tests is required to demonstrate SRO performance. In this paper, we report results of the on-beam validation of the Jefferson Lab SRO framework. In this work, we exposed different detectors (PbWO-based electromagnetic calorimeters and a plastic scintillator hodoscope) to the Hall-D electron-positron secondary beam and to the Hall-B production electron beam, with increasingly complex experimental conditions. By comparing the data collected with the SRO system against the traditional DAQ, we demonstrate that the SRO performs as expected. Furthermore, we provide evidence of its superiority in implementing sophisticated AI-supported algorithms for real-time data analysis and reconstruction.

47 OTHER INSTRUMENTATION↗

A causal data fusion method for the general exposure and outcome

Abstract With the advent of the big data era, the need to combine multiple individual data sets to draw causal effects arises naturally in many medical and biological applications. Especially each data set cannot measure enough confounders to infer the causal effect of an exposure on an outcome. In this article, we extend the method proposed by a previous study to causal data fusion of more than two data sets without external validation and to a more general (continuous or discrete) exposure and outcome. Theoretically, we obtain the condition for identifiability of exposure effects using multiple individual data sources for the continuous or discrete exposure and outcome. The simulation results show that our proposed causal data fusion method has unbiased causal effect estimate and higher precision than traditional regression, meta‐analysis and statistical matching methods. We further apply our method to study the causal effect of BMI on glucose level in individuals with diabetes by combining two data sets. Our method is essential for causal data fusion and provides important insights into the ongoing discourse on the empirical analysis of merging multiple individual data sources.

Li, Hongkai↗