Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel Programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Atomic-Scale Imaging Reveals Polar-π Interactions in Two-Dimensional Molecular Superlattices

Controlling coassembly of synthetic oligomers into binary superlattices at the atomic level is challenging. Here, we report a strategy for programming polar-π interactions in oligomeric peptoids, a class of sequence-defined peptidomimetics, facilitating the formation of homogeneous two-dimensional (2D) superlattices. N-2-phenylethyl and N-(2-perfluorophenyl)ethyl side chains, similar in size, but with contrasting electrostatic characteristics, were introduced at defined sequence positions to generate favorable dipolar aromatic interactions. The resulting nanosheets exhibit different crystal motifs depending on the side chain interactions: systems containing only one type of aromatic side chain form a parallel V-shaped motif driven by π-π interactions, whereas a combination of both types of aromatic side chains, either within one backbone or through the coassembly of two distinct peptoids, adopt an antiparallel V-shaped superlattice with higher thermal stability, driven by polar-π interactions. Cryogenic transmission electron microscopy directly resolved the packing arrangement of perfluorophenyl and phenyl rings in individual nanosheet superlattices, confirming that intermolecular polar-π interaction dominates the superlattice motifs and increases lattice stability. Molecular dynamics simulations and density functional theory calculations further substantiate the energetic favorability of polar-π interactions over π-π interactions, rationalizing the formation of homogeneous superlattices with enhanced thermal stability. Our discoveries establish a design principle for binary coassembly using sequence-defined oligomers, which enables control over unit cell geometry, lattice stability, and molecular registration through aromatic side chain polarization and sequence control. This ability to program atomic-scale binary superlattices opens new avenues for designing functional 2D soft materials.

Lee, Yen Jea [Lawrence Berkeley National Laborator↗

LHC Event Generation in the Exascale Era

MCFM is a dedicated Monte-Carlo simulation program for collider phenomenology at highest energies. Designed during the Tevatron era, it has successfully incorporated the latest developments needed for LHC precision calculations and remained on the forefront of collider phenomenology. The Fortran code includes interfaces to modern PDF and loop reduction libraries but has been unchanged structurally compared to the earlier versions. Parallel computing has been enabled using OpenMP and MPI. MCFM provides numerically highly stable one-loop amplitudes and superior phase-space efficiency, leading to excellent performance in NXLO calculations using jettiness or qT subtraction techniques for IR regularization.

Campbell, John [Fermilab]↗

HTGR Multiphysics Application Drivers FY26 Updates

This report summarizes FY26 progress under the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program's high-temperature gas-cooled reactor (HTGR) application driver work, covering a wide range of activities such as code validation and multi-physics code assessment. 1) A detailed SAM model of the High-Temperature Engineering Test Reactor (HTTR) was developed using a unique-block grouping approach, with an extended parallel thermal network method to capture block-to-block conduction and radiation heat transfer, and applied to steady-state simulations of the HTTR 30~MW and 9~MW cases. 2) In another activity, SAM's newly implemented multi-component gas flow model was validated against the Natural convection Shutdown heat removal Test Facility (NSTF) argon ingress experiment, correctly capturing the density-driven suppression and thermal recovery of natural circulation observed when argon is introduced into the air-cooled Reactor Cavity Cooling System (RCCS) loop. 3) For the OECD/NEA High Temperature Test Facility (HTTF) benchmark, we co-led the international benchmark activities as well as the OECD/NEA final benchmark report to be released at the end of this year. 4) Finally, the coupled Griffin-SAM modeling capability for pebble-bed HTGRs was advanced by verifying the Griffin neutronics solution against Serpent Monte Carlo for a realistic non-uniform temperature distribution, resolving several deficiencies in the SAM-to-Griffin temperature transfer scheme, and enabling distinct fuel kernel, moderator, and coolant temperatures for cross section feedback. These new features were demonstrated in a PBR load-following transient.

Lee, Alvin↗

Operational experience and R&D results using the Google Cloud for High-Energy Physics in the ATLAS experiment

The ATLAS experiment at CERN relies on a Worldwide Distributed Computing Grid infrastructure to support its physics program at the Large Hadron Collider. ATLAS has integrated cloud computing resources to complement its Grid infrastructure and conducted an R&D program on Google Cloud Platform. These initiatives leverage key features of commercial cloud providers: lightweight configuration and operation, elasticity and availability of diverse infrastructures. Here this paper examines the seamless integration of cloud computing services as a conventional Grid site within the ATLAS workflow management and data management systems, while also offering new setups for interactive, parallel analysis. It underscores pivotal results that enhance the on-site computing model and outlines several R&D projects that have benefited from large-scale, elastic resource provisioning models. Furthermore, this study discusses the impact of cloud-enabled R&D projects in three domains: accelerators and AI/ML, ARM CPUs and columnar data analysis techniques.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

ELECTRONIC STRUCTURE METHODS AND PROTOCOLS WITH APPLICATION TO DYNAMICS, KINETICS AND THERMOCHEMISTRY

Hydrocarbon combustion involves the reaction dynamics of a tremendous number of species beginning with many-component fuel mixtures and proceeding via a complex system of intermediates to form primary and secondary products. Combustion conditions corresponding to new advanced engines and/or alternative fuels rely increasingly on autoignition and low-temperature-combustion chemistry. In these regimes various transient radical species such as HO2, ROO·, ·QOOH, HCO, NO2, HOCO, and Criegee intermediates play important roles in determining the detailed as well as more general dynamics. A clear understanding and accurate representation of these processes is needed for effective modeling. Given the difficulties associated with making reliable experimental measurements of these systems, computation can play an important role in developing these energy technologies. Accurate calculations have their own challenges since even within the simplest dynamical approximations such as transition state theory, the rates depend exponentially on critical barrier heights and these may be sensitive to the level of quantum chemistry. Moreover, it is well-known that in many cases it is necessary to go beyond statistical theories and consider the dynamics. Quantum tunneling, resonances, radiative transitions, and non-adiabatic effects governed by spin-orbit or derivative coupling can be determining factors in those dynamics. Building upon progress made during a period of prior support through the DOE Early Career Program, this project combines developments in the areas of potential energy surface (PES) fitting and multistate multireference quantum chemistry to allow spectroscopically and dynamically/kinetically accurate investigations of key molecular systems (such as those mentioned above), many of which are radicals with strong multireference character and have the possibility of multiple electronic states contributing to the observed dynamics. An ongoing area of investigation is to develop general strategies for robustly convergent electronic structure theory for global multichannel reactive surfaces including diabatization of energy and other relevant surfaces such as dipole transition. Combining advances in ab initio methods with automated interpolative PES fitting allows the construction of high-quality PESs (incorporating thousands of high-level data) to be done rapidly through parallel processing on high-performance computing (HPC) clusters. In addition, new methods and approaches to electronic structure theory will be developed and tested through applications. This project will explore limitations in traditional multireference calculations (e.g., MRCI) such as those imposed by internal contraction, lack of high-order correlation treatment and poor scaling. Methods such as DMRG-based extended active-space CASSCF and various Quantum Monte Carlo (QMC) methods will be applied (including VMC/DMC and FCIQMC). Insight into the relative significance of different orbital spaces and the robustness of application of these approaches on leadership class computing architectures will be gained. Synergy with other components of this research program such as automated PES fitting and multireference quantum chemistry will be used to address challenges encountered by the standard approaches to computational thermochemistry (those being single-reference quantum chemistry and perturbative treatments of the anharmonic vibrational energy, which break down for some cases of electronic structure or floppy strongly coupled vibrational modes).

74 ATOMIC AND MOLECULAR PHYSICS↗

Alignment switching in 3D-printed smectic liquid crystal elastomers

Extrusion-based additive manufacturing has emerged as a powerful platform for designing shape-morphing materials through controlled orientation. However, existing approaches primarily rely on a single mode of flow-induced alignment, limiting orientation programmability. Herein, we present a direct-ink-writing approach for smectic liquid crystal elastics that exploits two distinct alignment modes within a single ink. The smectic ink exhibits shear- and temperature-dependent orientation switching, enabling molecular alignment either perpendicular or parallel to the print direction. Combined rheological, X-ray, and molecular dynamics analyses reveal that this alignment inversion arises from the preservation or collapse of smectic layers under flow. This reversible switching encodes both contractile and elongational actuation within individual filaments, greatly expanding the design freedom of printed liquid crystal elastomers. We demonstrate 2D and 3D structures with diverse programmed shape transformations, highlighting the potential of this platform for adaptive soft actuators and architected functional materials.

Lee, Jin Hyeong [Pusan National University, Busan,↗

The high level trigger and express data production at STAR

To meet the demands of the Beam Energy Scan phase-II (BES-II) program, the STAR experiment at the Relativistic Heavy Ion Collider (RHIC) developed a dual real-time framework consisting of a High Level Trigger (HLT) and an Express Data Production system (xProduction). The HLT operates online within the Data Acquisition (DAQ) chain on a dedicated multi-core CPU cluster with the option to offload compute-intensive kernels to Xeon Phi coprocessors. It uses parallelized algorithms, such as the Cellular Automaton (CA) Track Finder, to perform rapid tracking, vertexing, and event filtering. This allows it to select events of interest in real time and provide immediate feedback on detector and beam conditions. In contrast, the xProduction workflow runs concurrently and independently of the DAQ loop. It applies near offline-quality calibration and reconstruction within hours of data collection. The xProduction input is the express data stream, whose content can be enriched by HLT trigger/priority selections under DAQ/HLT resource constraints, and it uses the STAR calibration/conditions framework, incorporating online calibration/QA information when available. This enables early preliminary physics analysis, including the reconstruction of rare signals, such as hyperons and hypernuclei. It also provides collaboration-wide access to analysis-ready datasets. Together, the HLT and xProduction systems form a complementary architecture: the HLT performs online event selection while the xProduction chain delivers high-quality results within a short amount of time. This integrated framework has enabled the prompt reconstruction of the $^5_Λ$ He hypernucleus with high statistical significance and the efficient processing of hundreds of millions of heavy-ion collision events. In conclusion, its demonstrated scalability and robustness establish a model for future high-luminosity experiments requiring both online event filtering and rapid access to analysis-quality data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

MTUQ: a framework for estimating moment tensors, point forces, and their uncertainties

SUMMARY We introduce MTUQ, an open-source Python package for seismic source estimation and uncertainty quantification, emphasizing flexibility and operational scalability. MTUQ provides MPI-parallelized grid search and global optimization capabilities, compatibility with 1-D and 3-D Green’s function database formats, customizable data processing, C-accelerated waveform and first-motion polarity misfit functions, and utilities for plotting seismic waveforms and visualizing misfit and likelihood surfaces. Applicability to a range of full- and constrained-moment tensor, point force, and centroid inversion problems is possible via a documented application programming interface, accompanied by example scripts and integration tests. We demonstrate the software using three different types of seismic events: (1) a 2009 intraslab earthquake near Anchorage, Alaska; (2) an episode of the 2021 Barry Arm landslide in Alaska; and (3) the 2017 Democratic People’s Republic of Korea underground nuclear test. With these events, we illustrate the well-known complementary character of body waves, surface waves, and polarities for constraining source parameters. We also convey the distinct misfit patterns that arise from each individual data type, the importance of uncertainty quantification for detecting multimodal or otherwise poorly constrained solutions, and the software’s flexible, modular design.

58 GEOSCIENCES↗

Evaluation of flow-induced plate deflection for University of Missouri research reactor low-enriched uranium fuel element

The University of Missouri Research Reactor (MURR), located on the campus of the University of Missouri in Columbia, Missouri, is one of the six United States (U.S.) High Performance Research Reactors (USHPRR), including one critical facility, that are actively collaborating with the U.S. Department of Energy (DOE) National Nuclear Security Administration (NNSA) Office of Material Management and Minimization (M3) Reactor Conversion Program to convert from highly enriched uranium (HEU, ≥20 wt% U-235) fuel to low-enriched uranium (LEU, <20 wt% U-235) fuel. A new type of very high-density LEU fuel based on a monolithic alloy of uranium and 10 wt% molybdenum (U-10Mo) is expected to allow conversion of some USHPRR, including MURR. In the design of its fuel elements, MURR is using thin parallel curved fuel plates separated by coolant channels. In this work, fluid-structure interaction (FSI) analysis of the MURR LEU fuel element is performed at the element level (as compared to the plate level analysis), which models all components of the LEU fuel element, including fuel plates and the supporting structures. Therefore, the effect of supporting structures on the flow distribution within the element and the fuel plate deflection are evaluated. In addition to the element nominal flow rate and dimensions, the tolerances in the geometry of the coolant channel and plate thickness, the effect of a comb on plate deflection, and the uncertainty of the flow rate per element are evaluated. For the LEU fuel plates, which are thinner than the current HEU plates, the predicted plate deflection is found to be small compared to the fabrication and assembly tolerances. Thus, the FSI-induced deflections are not expected to noticeably reduce the coolant flow rate or predicted safety margins in the limiting channels for the MURR LEU fuel element. In addition to the simulation work, a hydraulic performance test of the MURR LEU fuel element is currently being planned to support conversion to the use of LEU fuel.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Identification and mitigation of memory block timing issue in ITk ABCStar during ASIC production

The ABCStar is a mixed-signal front-end readout ASIC for the strips sensor portion of the ATLAS ITk detector being developed as part of the High-Luminosity LHC upgrade. In pre-production testing, a subtle design flaw was uncovered in the ABCStar that was reducing wafer yields in some manufactured lots from the expected 90% to as low as 2%. The root cause was determined to be a timing issue in the logic synthesized to control previously silicon proven memory blocks re-used for this ASIC. The solutions proposed included manufacturing process changes by the wafer foundry, changes to the operating parameters for the ABCStar in the detector, and the possibility that a redesign might be required. The two mitigation efforts were undertaken in parallel, with the process modification route a less desirable solution since already manufactured wafers would need to be scrapped in favour of the new ones. Based on a knowledge of the existing process, and testing done on the worst performing wafers, it was proposed that raising the core operating voltage of the ABCStar from 1.20V to 1.25V could address the timing issue by sufficiently speeding up its transistors. An extensive testing program that included the effects of temperature and radiation expected over the lifetime of the ITk detector was conducted to validate that approach. Those tests and studies proved that even the worst performing wafers would have yields over 80% with the 1.25V core voltage, and neither the modified process nor redesign would be required for ensuring reliable operation of the ITk. Based on testing, a further timing mitigation was implemented to provide an additional margin of reliability by increasing the duty cycle of the clock to the ABCStar. Testing of all ABCStar wafers has been completed and the production of the detector modules using these ASICs is now well underway as a result of the efforts detailed herein.

FOS: Physical sciences↗

T RI M E ++: Multi-threaded triangular meshing in two dimensions

We present T RI M E ++, a multi-threaded software library designed for generating two-dimensional meshes for intricate geometric shapes using the Delaunay triangulation. Multi-threaded parallel computing is implemented throughout the meshing procedure, making it suitable for fast generation of large-scale meshes. Three iterative meshing algorithms are implemented: the DistMesh algorithm, the centroidal Voronoi diagram meshing, and a hybrid of the two. We compare the performance of the three meshing methods in T RI M E ++, and show that the hybrid method retains the advantages of the other two. The software library achieves significant parallel speedup when generating large-scale meshes containing between 10 4 to 10 7 points. T RI M E ++ can handle complicated geometries and generates adaptive meshes of high quality.

97 MATHEMATICS AND COMPUTING↗

Leveraging FPGA Advantages for Quicker Data Processing for LBNF

The Long Baseline Neutrino Facility (LBNF) will deliver a 2.4 MW muon neutrino beam from Fermilab to the Deep Underground Neutrino Experiment (DUNE), requiring unprecedented precision in beamline alignment to achieve DUNE's neutrino oscillation measurement goals. Vertical misalignments of beamline components as small as 0.5 mm can contribute 6-7\% uncertainty in predicted neutrino flux, necessitating sub-0.1 mm alignment monitoring capabilities. The Horn Location Sensor (HLS) system employs frequency sweep interferometry (FSI) in a distributed hydrostatic leveling network to achieve the required precision under harsh radiation conditions up to 5000 kRad/year. Traditional FSI implementations suffer from laser sweep nonlinearities that degrade resolution and require computationally intensive post-processing corrections using gas reference cells. This work presents a real-time FPGA-based implementation of the HLS data acquisition and processing system using a sweep tracker interferometer for dynamic sweep linearization. The system utilizes a PYNQ-Z2 FPGA with programmable logic implementing parallel 16k-point FFT processing across four channels, synchronized by the sweep tracker signal to eliminate post-processing requirements. Spectral performance testing demonstrates significant improvements in peak sharpness compared to traditional fixed-frequency digitization. The FPGA implementation enables real-time displacement monitoring with processing speeds orders of magnitude faster than software-based approaches, essential for the operational requirements of LBNF's eventual distributed sensor network. This advancement in real-time FSI processing directly supports DUNE's precision neutrino physics program by providing the rapid feedback necessary for maintaining stringent beamline alignment tolerances during high-power beam operations.

Rossel, Jacob↗

Real-Time FPGA Implementation For Frequency Sweep Interferometry In The LBNF Complex

The Long Baseline Neutrino Facility (LBNF) will deliver a 2.4 MW muon neutrino beam from Fermilab to the Deep Underground Neutrino Experiment (DUNE), requiring unprecedented precision in beamline alignment to achieve DUNE's neutrino oscillation measurement goals. Vertical misalignments of beamline components as small as 0.5 mm can contribute 6-7\% uncertainty in predicted neutrino flux, necessitating sub-0.1 mm alignment monitoring capabilities. The Horn Location Sensor (HLS) system employs frequency sweep interferometry (FSI) in a distributed hydrostatic leveling network to achieve the required precision under harsh radiation conditions up to 5000 kRad/year. Traditional FSI implementations suffer from laser sweep nonlinearities that degrade resolution and require computationally intensive post-processing corrections using gas reference cells. This work presents a real-time FPGA-based implementation of the HLS data acquisition and processing system using a sweep tracker interferometer for dynamic sweep linearization. The system utilizes a PYNQ-Z2 FPGA with programmable logic implementing parallel 16k-point FFT processing across four channels, synchronized by the sweep tracker signal to eliminate post-processing requirements. Spectral performance testing demonstrates significant improvements in peak sharpness compared to traditional fixed-frequency digitization. The FPGA implementation enables real-time displacement monitoring with processing speeds orders of magnitude faster than software-based approaches, essential for the operational requirements of LBNF's eventual distributed sensor network. This advancement in real-time FSI processing directly supports DUNE's precision neutrino physics program by providing the rapid feedback necessary for maintaining stringent beamline alignment tolerances during high-power beam operations.

Rossel, A. Jacob [Fermilab; Unlisted]↗

ACHILLES-GENIE Interface for Neutrino Simulations, ADRIANO2 Tile Prototype for High-Granularity Dual-Readout Calorimetry

ICARUS (Imaging Cosmic And Rare Underground Signals) is a liquid argon time projection chamber (LArTPC) detector that pursues the sterile neutrino, which relies on accurate simulations of neutrino-argon interactions. REDTOP (Rare Eta Decays To Observe new Physics) is a proposed low-energy, high-intensity meson factory designed to explore rare $\eta$/$\eta'$ meson decays and probe physics beyond the Standard Model. As a next-generation experiment, this requires both accurate simulations and innovative detector technologies. This project contributes to both ICARUS, from a simulation perspective, and REDTOP, from both a simulation and detection perspective, through the event generation of lepton-nucleon interactions and the physical enhancement of the calorimeter technology within the REDTOP detector. We developed an interface between ACHILLES (A CHIcago Land Lepton Event Simulator), a theory-driven lepton-level event generator, and GENIE, a robust event generator framework used for neutrino physics. By incorporating the precise theoretical cross-section calculations of ACHILLES into the experimental realism of GENIE, the interface allows for improved accuracy of neutrino-nucleon simulations, which can be adapted for the proton beam specifications of the REDTOP meson factory as well as for the ICARUS experiment. In parallel, we developed an improved prototype for the ADRIANO2 (A Dual Readout Integrally Active Non-segmented Option) dual-readout calorimeter tiles for the REDTOP detector. To improve the efficiency of the lead-glass tiles trapping Cherenkov light for energy reconstruction and particle identification, we optimized the application of a highly reflective coating. Through viscosity and thickness control, masking, and a custom spray technique, we refined the coating process to reduce surface defects and improve light yield. Together, these efforts strengthen the ICARUS neutrino program and REDTOP's capability of detecting rare decay events.

Visser, Erin [Michigan State U.]↗

Enhancing ICARUS and REDTOP Software and Hardware: Event Generator Interface Development and Calorimeter Tile Prototype

ICARUS (Imaging Cosmic And Rare Underground Signals) is a liquid argon time projection chamber (LArTPC) detector that pursues the sterile neutrino, which relies on accurate simulations of neutrino-argon interactions. REDTOP (Rare Eta Decays To Observe new Physics) is a proposed low-energy, high-intensity meson factory designed to explore rare $\eta$/$\eta'$ meson decays and probe physics beyond the Standard Model. As a next-generation experiment, this requires both accurate simulations and innovative detector technologies. This project contributes to both ICARUS, from a simulation perspective, and REDTOP, from both a simulation and detection perspective, through the event generation of lepton-nucleon interactions and the physical enhancement of the calorimeter technology within the REDTOP detector. We developed an interface between ACHILLES (A CHIcago Land Lepton Event Simulator), a theory-driven lepton-level event generator, and GENIE, a robust event generator framework used for neutrino physics. By incorporating the precise theoretical cross-section calculations of ACHILLES into the experimental realism of GENIE, the interface allows for improved accuracy of neutrino-nucleon simulations, which can be adapted for the proton beam specifications of the REDTOP meson factory as well as for the ICARUS experiment. In parallel, we developed an improved prototype for the ADRIANO2 (A Dual Readout Integrally Active Non-segmented Option) dual-readout calorimeter tiles for the REDTOP detector. To improve the efficiency of the lead-glass tiles trapping Cherenkov light for energy reconstruction and particle identification, we optimized the application of a highly reflective coating. Through viscosity and thickness control, masking, and a custom spray technique, we refined the coating process to reduce surface defects and improve light yield. Together, these efforts strengthen the ICARUS neutrino program and REDTOP's capability of detecting rare decay events.

Visser, Erin [Michigan State U.] (ORCID:0009000184↗

Direct NeTS sampling of nuclear graphite $S(α, β, T)$ in Serpent

For advanced reactor applications, Neural Thermal Scattering (NeTS) modules were developed to predict the thermal scattering law (TSL or $S(α, β, T)$) of a nuclear graphite neutron moderator. NeTS are multi-layer, feedforward artificial neural networks, which act as universal function approximators designed for TSL datasets. In this case, a 4-layer neural network with 164 neurons per layer is trained using FLASSH evaluated data in PyTorch and serialized as a torchscript dictionary to predict $S(α, β, T)$ on-the-fly. Relative, absolute and maximum percent deviations of NeTS from File 7 data generated using the FLASSH code are on the order of 0.01%, 0.1% and 1%, respectively, with low inference latencies of 0.000172 s per $S(α, β, T)$ at a given temperature. Capturing the full dimensionality of possible inelastic neutron-lattice interactions, NeTS functionality is embedded in the Serpent Monte Carlo code, where $S(α, β, T)_{NeTS}$ sampling is conducted on-the-fly and compared to ACE look-up-tables for predicting TREAT criticality. k-eff differences between sampling algorithms of 6 pcm are observed and are within the order of Monte Carlo uncertainty. Compared to discrete and continuous-energy ACE files (30 MB and 131 MB per temperature), the NeTS format is on the order of 200–300 kB for a continuous-temperature, interpolation-free representation of $S(α, β, T)$ and cross sections. NeTS-in-Serpent runtimes comparable with ACE look-up tables are achieved by scaling NeTS for high performance computing architectures with hybrid OpenMP + MPI parallelization. This work validates a novel, self-contained reactor physics framework for predictive cross sections, and demonstrates a general methodology for embedding modern machine learning libraries within existing neutronic analysis frameworks.

Nuclear Criticality Safety Program (NCSP)↗

SPARTAN (Scalable Probabilistic Application Reconfigurable Tensor Autonomous Network)

The technical founder of Ludwig Computing Inc has been competitively selected for support by Cyclotron Road, a U.S. Department of Energy (DOE) Advanced Manufacturing Office (AMO) Lab-Embedded Entrepreneurship Program (LEEP) through an approved merit review process. Ludwig Computing Inc, supported by the U.S. Department of Energy's Advanced Manufacturing Office through the Cyclotron Road program, has investigated the advantages of probabilistic computing for real-world compute-intensive applications. This research adds to the understanding of alternative computing paradigms by exploring a unique hardware-software co-design that integrates quantum computing methods with nature-inspired problem-solving techniques. The project's focus on areas such as combinatorial optimization, graph analytics, and machine learning demonstrates the potential for significant advancements in computational efficiency and performance. By harnessing natural randomness to streamline large circuits into fewer devices, Ludwig's approach enables massive parallelism, potentially offering higher throughput, speed, and energy efficiency compared to conventional hardware solutions. This work benefits the public by paving the way for more efficient computing solutions that could address complex real-world problems while potentially reducing energy consumption in data-intensive industries.

97 MATHEMATICS AND COMPUTING↗

Accelerating Neutrino Event Generation in MARLEY Using CUDA-Based RNG and GPU Parallelization

MARLEY is a simulation tool that helps scientists study how low-energy neutrinos interact with matter. To work properly, MARLEY uses random numbers thousands of times in each simulation. These random numbers are important for modeling things like how neutrinos collide with atoms and what particles they produce. Right now, MARLEY runs on a regular computer processor (CPU) and uses a built-in random number generator called the Mersenne Twister. This setup works, but it can be slow, especially when trying to simulate many events. This research focuses on making MARLEY run faster by moving the random number generation and some of the repetitive calculations from the CPU to a graphics processing unit (GPU), which can handle many tasks at the same time. We use CUDA (a tool for programming NVIDIA GPUs) and cuRAND (a GPU-based random number library) to test faster alternatives to the current random number system. We compare different GPU-based generators, like curand_mtgp32, xorwow, and philox, to see which ones are the quickest and still give reliable results. Early tests show that using the GPU can make MARLEY simulations much faster. This project not only helps improve current simulation performance but also moves closer to a full simulation chain where all stages can run on modern GPU hardware.

Dunkley, Kimieka [Florida A-M]↗