Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “NiC”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Preparing MPICH for exascale

The advent of exascale supercomputers heralds a new era of scientific discovery, yet it introduces significant architectural challenges that must be overcome for MPI applications to fully exploit its potential. Among these challenges is the adoption of heterogeneous architectures, particularly the integration of GPUs to accelerate computation. Additionally, the complexity of multithreaded programming models has also become a critical factor in achieving performance at scale. The efficient utilization of hardware acceleration for communication, provided by modern NICs, is also essential for achieving low latency and high throughput communication in such complex systems. In response to these challenges, the MPICH library, a high-performance and widely used Message Passing Interface (MPI) implementation, has undergone significant enhancements. Here, this paper presents four major contributions that prepare MPICH for the exascale transition. First, we describe a lightweight communication stack that leverages the advanced features of modern NICs to maximize hardware acceleration. Second, our work showcases a highly scalable multithreaded communication model that addresses the complexities of concurrent environments. Third, we introduce GPU-aware communication capabilities that optimize data movement in GPU-integrated systems. Finally, we present a new datatype engine aimed at accelerating the use of MPI derived datatypes on GPUs. These improvements in the MPICH library not only address the immediate needs of exascale computing architectures but also set a foundation for exploiting future innovations in high-performance computing. By embracing these new designs and approaches, MPICH-derived libraries from HPE Cray and Intel were able to achieve real exascale performance on OLCF Frontier and ALCF Aurora respectively.

Guo, Yanfei [Argonne National Laboratory (ANL), Ar↗

Comparison of All Solid Cancer Mortality and Incidence Dose-Response in the Life Span Study of Atomic Bomb Survivors, 1958–2009

Recent analysis of all solid cancer incidence (1958–2009) in the Life Span Study (LSS) revealed evidence of upward curvature in the radiation dose response among males but not females. Upward curvature in sex-averaged excess relative risk (ERR) for all solid cancer mortality (1950–2003) was also observed in the 0–2 Gy dose range. As reasons for non-linearity in the LSS are not completely understood, we conducted dose response analyses for all solid cancer mortality and incidence applying similar methods (1958–2009 follow-up, DS02R1 doses, including subjects notin-city (NIC) at the time of the bombing) and statistical models. Incident cancers were ascertained from Hiroshima and Nagasaki cancer registries, while cause of death was ascertained from death certificates over entire Japan. The study included 105,444 LSS subjects who were alive and not known to have cancer before Jan 1, 1958 (80,205 with dose estimates and 25,239 NIC subjects). Between 1958 and 2009, there were 3.1 million person-years (PY) and 22,538 solid cancers for incidence analysis and 3.8 million PY and 15,419 solid cancer deaths for mortality analysis. We fitted sex-specific ERR models adjusted for smoking to both types of data. Over the entire range of doses, solid cancer mortality dose response exhibited a borderline significant upward curvature among males (P=0.062) and significant upward curvature among females (P=0.010); for solid cancer incidence, as before, we found a significant upward curvature among males (P=0.001) but not among females (P=0.624). The sex difference in magnitude of dose response curvature was statistically significant for cancer incidence (P=0.017) but not for cancer mortality (P=0.781). The results of analyses in the 0–2 Gy range and restricted lower dose ranges generally supported inferences made about the sex-specific dose response shape over the entire range of doses for each outcome. Patterns of sex-specific curvature by calendar period (1958–1987 vs 1988–2009) and age at exposure (0–19 vs 20–83) varied between mortality and incidence data, particularly among females, although for each outcome there was an indication of curvature among 0–19 year old male survivors in both calendar periods and among 0–19 year old female survivors in the recent period. Collectively, our findings indicate that the upward curvature in all solid cancer dose response in the LSS is neither specific to males nor to incidence data; it appears to depend on composition of case series and age at exposure or time. Further follow-up and site-specific analyses of cancer mortality and incidence will be important to confirm the emerging trend in dose response curvature among young survivors and unveil the contributing factors and sites.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Substituent effect on napthodithiophene-fused porphyrins: Understanding the unusual trend of fluorescence quantum yield quenching

A series of π-extended porphyrins containing thiophene units was prepared to investigate an unusual trend observed in the fluorescence quantum yield of naphtho[2,1-b:3,4-b']dithiophene-fused porphyrins. Monobenzoporphyrins carrying 2-thiophenyl (vinyl thiophene porphyrin) groups (2VTP, 2VTBr and Br2VTP) (numbering 2 and 3 refers to the position of sulfur on the thiophene ring) were synthesized through a Heck-based cascade reaction followed by a newly developed bromination method. These were converted to naphtho[2,1-b:3,4-b']dithiophene-fused porphyrin derivatives (FBr2VTP and F2VTBr) via ring closure with an intramolecular Scholl reaction. Increasingly larger Stokes shifts with a higher number of bromo groups were observed in the unfused π-extended molecular systems, reflecting the heavy atom effect. Fluorescence spectroscopy further confirmed the unusual trend seen in the previous work: structural rigidification in naphtho[2,1-b:3,4-b']dithiophene-fused porphyrins leads to longer fluorescence lifetimes but unexpectedly lowers the quantum yield. Adding one bromo group to the naphtho[2,1-b:3,4-b']dithiophene unit does not change this trend. Furthermore, the presence of two bromo groups prevents the quantum yield from dropping. DFT, TDDFT, and NICS analyses suggest a drastic change in aromaticity in the pyrrole ring where the naphtho[2,1-b:3,4-b']dithiophene is fused in FBr2VTP, which might contribute to the quantum yield reductions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Probing Effects of Electron Addition to Polycyclic Conjugated Hydrocarbons Containing Four‐Membered Rings

The rectangular cyclobutadiene (CBD, C 4 H 4 ) is a unique moiety for building nonbenzenoid polycyclic conjugated hydrocarbons with interesting electron‐accepting properties. Herein, the investigation on chemical reduction of several CBD‐containing polycyclic hydrocarbons with increasing conjugation length is reported: biphenylene (C 12 H 8 ), dimethyl[2]naphthalene (C 22 H 16 ), and tetramethyl‐dibenzo‐[3]phenylene (C 30 H 22 ). The two‐step sequential reduction is first demonstrated by in situ spectroscopic investigation and then confirmed by the isolation of single crystals of the reduced products. The X‐ray crystallographic analysis reveals the formation of several mono‐ and doubly reduced products in solvent‐separated and complexed forms. The crystal structures for both neutral parents and corresponding reduced products unravel the changes in bond alternation in each ring of the fused systems. Density functional theory (DFT) and nucleus‐independent chemical shift (NICS) scan calculations reveal that the two‐electron addition reduces the aromatic character in the benzenoid rings but has minor influence on the antiaromatic CBD rings.

Zhu, Yikun [Department of Chemistry University at ↗

Dual active site tandem catalysis of metal hydroxyl oxides and single atoms for boosting oxygen evolution reaction

We report the high voltage in oxygen evolution reaction (OER) often causes structural change in electrocatalysts and forms multiple active sites. Therefore, exploring the synergy of various active sites is extremely significant to develop catalysts for OER with multiple elementary steps. Herein, we adopt the physically adsorbed metal ions method to successfully synthesize a highly-efficient electrocatalyst containing the dual active sites of Fe-NiOOH and NiC4 single atoms (marked as Ni SAs/Fe-NiOOH). The Ni SAs/Fe-NiOOH displays outstanding OER performance with an overpotential of 269 mV to deliver current density of 10 mA/cm 2 that shows 55 mV superior to commercial IrO 2 /CB at the same condition. Experiments and density functional theory calculations indicate that the excellent OER activity of Ni SAs/Fe-NiOOH catalyst is attributed to the synergy of dual active sites of NiC 4 SAs and Fe-NiOOH. A tandem catalysis mechanism is also proposed to reveal the synergism of two active centers, which makes the potential-determining step more facile and accordingly decreases the OER overpotential. This work offers a new concept of tandem catalysis to develop the electrocatalysts with many elemental steps, like OER and oxygen reduction reaction catalysts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Optically connected memory for disaggregated data centers

Recent advances in integrated photonics enable the implementation of reconfigurable, high-bandwidth, and low energy-per-bit interconnects in next-generation data centers. We propose and evaluate an Optically Connected Memory (OCM) architecture that disaggregates the main memory from the computation nodes in data centers. OCM is based on micro-ring resonators (MRRs), and it does not require any modification to the DRAM memory modules. We calculate energy consumption from real photonic devices and integrate them into a system simulator to evaluate performance. Here, our results show that (1) OCM is capable of interconnecting four DDR4 memory channels to a computing node using two fibers with 1.02 pJ energy-per-bit consumption and (2) OCM performs up to 5.5× faster than a disaggregated memory with 40G PCIe NIC connectors to computing nodes.

97 MATHEMATICS AND COMPUTING↗

Affinity of Nicotinoids to a Model Nicotinic Acetylcholine Receptor (nAChR) Binding Pocket in the Human Brain

The binding affinity of nicotinoids to the binding residues of the a4ß2 variant of the nicotinic acetylcholine receptor (nAChR) was identified as a strong predictor of the nicotinoid’s addictive character. Using ab-iniito calculations for model binding pockets of increasing size comprising of 3, 6, and 14 amino acids (3AA, 6AA, and 14AA) that are derived from the crystal structure, the differences in binding affinity of 6 nicotinoids, namely nicotine (NIC), nornicotine (NOR), anabasine (ANB), anatabine (ANT), myosmine (MYO), and cotinine (COT) were correlated to their previously reported doses required for increases in intracranial self-stimulation (ICSS) thresholds, a metric for their addictive function. By employing the many body decomposition, the differences in the binding affinities of the various nicotinoids could be attributed mainly to the proton exchange energy between the Pyridine and non-Pyridine rings of the nicotinoids and the interactions between them and a handful of proximal amino acids, namely Trp156, Trpß57, Tyr100, and Tyr204. Interactions between the guest nicotinoid and the amino acids of the binding pocket were found to be mainly classical in nature, except for those between the nicotinoid and Trp156. The larger pockets were found to model binding structures more accurately and predicted the addictive character of all nicotinoids while smaller models, which are more computationally feasible, would only predict the addictive character of nicotinoids that are similar to nicotine. Here, the present study identifies the binding affinity of the guest nicotinoid to the host binding pocket as a strong descriptor of the nicotinoid’s addiction potential and as such it can be employed as a fast screening technique for the potential addiction of nicotine analogs.

60 APPLIED LIFE SCIENCES↗

HPC Network Simulation Tuning via Automatic Extraction of Hardware Parameters

Popular HPC network interconnection simulators such as SST/macro provide a variety of configurable parameters to explore the design space of hardware components such as network interface cards (NIC), switches, and links among them. While such knobs provide flexibility to explore design trade-offs for novel hardware, manually configuring simulations for matching configurations of the existing hardware to focus on topology exploration can be cumbersome and error-prone, leading to widely inaccurate simulations. This challenge is compounded when specifications of various (proprietary) technologies are not readily available or intentionally omitted. In this work, we propose a framework to autotune the multiple network models’ simulation configurations within SST/macro using Tree-structured Parzen Estimator-based Bayesian optimization to observe the effect on simulation accuracy across different message regimes. These regimes consist of small to large message sizes and latency to bandwidth-bound messages. We provide a detailed analysis of the simulation error for four representative HPC systems. Our Bayesian optimization based autotuning framework for network models achieves a maximum of 5x improvement in accuracy over best-effort manual configurations based on available hardware specifications.

Simulation, autotuning↗

A Case For Intra-rack Resource Disaggregation in HPC

The expected halt of traditional technology scaling is motivating increased heterogeneity in high-performance computing (HPC) systems with the emergence of numerous specialized accelerators. As heterogeneity increases, so does the risk of underutilizing expensive hardware resources if we preserve today’s rigid node configuration and reservation strategies. This has sparked interest in resource disaggregation to enable finer-grain allocation of hardware resources to applications. However, there is currently no data-driven study of what range of disaggregation is appropriate in HPC. To that end, we perform a detailed analysis of key metrics sampled in NERSC’s Cori, a production HPC system that executes a diverse open-science HPC workload. In addition, we profile a variety of deep-learning applications to represent an emerging workload. We show that for a rack (cabinet) configuration and applications similar to Cori, a central processing unit with intra-rack disaggregation has a 99.5% probability to find all resources it requires inside its rack. In addition, ideal intra-rack resource disaggregation in Cori could reduce memory and NIC resources by 5.36% to 69.01% and still satisfy the worst-case average rack utilization.

97 MATHEMATICS AND COMPUTING↗

eCounter: Inline Per-IP Network Monitoring at Millisecond Resolution via eBPF

Scientific data acquisition (SciDAQ) systems are shifting from archive-based workflows to streaming paradigms, where real-time, fine-grained network monitoring becomes essential. While P4-enabled devices offer per-packet in-band observability, they require specialized switches and routers. Host-side tools like Prometheus exporters lack sufficient temporal granularity. To bridge this gap, we present eCounter, a lightweight, hardware-agnostic, inline telemetry agent built on extended Berkeley Packet Filter (eBPF). eCounter captures per-interface ingress and egress traffic, categorized by IP address and protocol, at millisecond to sub-millisecond resolution. In a 100 Gbps environment, it continuously exports up to 3,257 time-series bins per second with only 4% CPU utilization at a 35¿KiB/s data rate. We evaluate eCounter across diverse NIC MTU settings, hook types, CPU architectures and operating systems, and observed negligible impact on concurrent high-throughput streaming applications. Complexity analysis confirms that it can be readily scaled to distributed SciDAQ deployments.

Mei, Xinxin [Computational Sciences and Technology↗

Frontier (HPE Cray EX) Exascale Supercomputer at the Oak Ridge Leadership Computing Facility

Frontier is the HPE Cray EX exascale supercomputer deployed and operated by the Oak Ridge Leadership Computing Facility (OLCF) at Oak Ridge National Laboratory (ORNL). Frontier is designed for large-scale modeling, simulation, and AI workloads and is built from HPE Cray EX system architecture with AMD CPUs and AMD Instinct GPU accelerators connected by the HPE Slingshot interconnect. System composition (representative production configuration): Frontier is composed of approximately 74 cabinets with 128 compute nodes per cabinet (~9,400 compute nodes total). Each compute node contains one 64-core AMD EPYC CPU and four AMD Instinct MI250X GPUs. Nodes are connected using HPE Slingshot (Slingshot-200 class) networking with multiple NIC ports per node providing high injection bandwidth. Frontier is connected to the Orion parallel file system (multi-tier Lustre) providing a large, center-wide high-performance storage namespace. Operational context: Frontier entered public prominence as the first system to reach No. 1 on the TOP500 list in May 2022 (HPL benchmark), establishing the first widely recognized exascale-era performance milestone. The system supports DOE Office of Science mission workloads and enables leadership-class computational science and AI for open science users.

AMD EPYC↗

Investigating Scientific Workload Acceleration using BlueField SmartNICs [Slides]

Modern computing platforms whose workloads generate large amounts of network traffic, such as cloud and HPC systems, often suffer from performance bottlenecks associated with the network interface. In order to alleviate the effects of this obstacle, a new generation of accelerators known as ‘SmartNICs’, which are designed to offload low level networking tasks from the processor into the NIC, have emerged.

42 ENGINEERING↗

Optimism is not a strategy: A white paper on how to give IFE a fighting chance to be real

With NIF shot N210808, we now have an existence proof of ignition (i.e. Lawson-like criteria exceeded and capsule gain well exceeding unity) in the laboratory and it has generated renewed interest in IFE. However, it is important to recognize that ignition on the NIF has been much more difficult than what was originally envisioned. Moreover, the design for the target that actually obtained burning plasma (Kritcher, Young, Robey, et al., Nature Phys. 2022; Zylstra, Hurricane, Callahan, et al., Nature, 601, 542, 2022) and ignition conditions is much different than the high gain design originally planned in the National Ignition Campaign (NIC; e.g. Lindl, Phys. Plasmas, 2, 3933, 1995; Lindl, Amendt, Berger, et al., Phys. Plasmas, 11, 339, 2004). In order to avoid squandering time and resources, the IFE community must learn the lessons of what happened on the NIF over the past decade.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Community Resilience Indicator Analysis: Commonly Used Indicators from Peer-Reviewed Research (Updated for Research Published 2003-2021)

In 2017, FEMA’s National Integration Center (NIC) Technical Assistance (TA) Branch identified a need to establish a data-driven basis for prioritizing locations for TA investment and guiding local emergency management planning. To achieve this goal, FEMA tasked Argonne National Laboratory (Argonne) with identifying commonly used indicators of community resilience across the landscape of published peer-reviewed research. FEMA and Argonne completed the first Community Resilience Indicator Analysis (CRIA) in 2018 and repeated the process in 2022. The CRIA process begins with a literature review and cataloguing of published peer-reviewed assessment methodologies on social vulnerability and community resilience. The literature review findings are then filtered by inclusion criteria established by the CRIA research team to ensure the methodologies are: (1) Quantitative, (2) Data and methodology are publicly available, (3) Calculated at the county level or lower, (4) Examine generalized hazard risk (rather than a singular hazard), and (5) Focused on pre-disaster community conditions. After this, the research team identifies the commonly used indicators across these methodologies and selects the best data source for each indicator. Finally, the research team bins the data for visual display, conducts a correlation analysis and creates a composite index, the FEMA Community Resilience Index (FEMA CRI). In 2018, the CRIA identified eight resilience and vulnerability assessment methodologies and 20 commonly used indicators (indicators used in three or more of the eight methodologies). The FEMA CRI in 2018 was created from these 20 indicators and was produced for at the county level. The 2022 CRIA updated the literature review to expand the list of methodologies examined and followed the same process, resulting in an analysis of 14 methodologies published between 2003 and 2021 and 22 indicators identified as commonly used (indicators used in five or more of the 14 methodologies). In 2022, the research team produced the FEMA CRI at the county and the census tract levels. To make the CRIA data more accessible and more actionable, each individual indicator and the FEMA CRI is binned and included in FEMA’s Resilience Analysis and Planning Tool (RAPT). RAPT enables emergency managers and community partners to quickly visualize relative differences in potential resilience by county, tribe and census tract. By reviewing the data for each of these 22 indicators individually, emergency managers can gain insights for targeted outreach strategies, planning, mitigation investments and response and recovery operations. Communities, regional governments and others can use this data to better understand potential challenges to resilience. As the social science field of examining and validating indicators of resilience evolves, FEMA will update RAPT to provide emergency managers and community partners with additional data and tools to inform planning, mitigation, response and recovery. It is important to understand that the role of the emergency manager is not to change or to “improve” the data, but to plan appropriately for the community characteristics reflected in the data. These datasets are community characteristics that researchers have identified as important considerations for resilience. For example, people with disabilities may have greater challenges to be resilient to disasters. If a community has a high population of people with disabilities, the emergency manager(s) may need to create tailored preparedness outreach programs and strategies to ensure those residents have support if evacuation is necessary. Rather than label these indicators as an absolute measure of resilience, FEMA considers “potential challenges to resilience” a better frame to understand these indicators. Everyone is vulnerable to disasters. While scholars theorize that certain characteristics may make an individual or a household more socially vulnerable, the data does not reflect measures that individuals and/or communities have taken to address potential challenges, such as emergency management planning and outreach or household preparedness measures. To aid emergency managers in understanding how to use these indicators, calling them potential challenges to resilience supports a more positive and strategic application of the data in all phases of emergency management.

99 GENERAL AND MISCELLANEOUS↗

Why are we still here 80 years later?

The Lab helped end World War II in 1945, but could the Lab survive peace? Join Senior Lab Historian Alan Carr and Historian Nic Lewis as they present “Why we’re still in Business,” examining the role of the Lab once World War II had ended. This presentation, which commemorates the Lab’s 80th anniversary, Carr and Lewis reconsider the popular belief that the Lab nearly ceased to exist after having developed the atomic bombs that helped end war.

99 GENERAL AND MISCELLANEOUS↗

In-Network DAQ Functions

A revolution in networking is changing how we compute, but we lack the tools that can channel this new capability to benefit science. It is now possible to write programs that operate "in" the network—on network cards (NICs) and network switches themselves, rather than on servers. These programs can analyze and reduce huge volumes of data as they flow through the network—at higher throughput, lower latency, and lower power consumption than if servers (containing CPUs or GPUs) were used. That equipment offers appealing features for scientific experiments that involve huge quantities of data. This poster describes a prototype LArTPC raw-waveform hit finder based on DUNE’s Trigger Primitives generator. We built this as part of ongoing research to better understand how to put programmable network hardware to use in large scientific experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

LEED: A Lightwave Energy-Efficient Datacenter

The Lightwave Energy-Efficient Datacenter (LEED) program is a disruptive “green-field” approach that provides a quantum leap in the energy efficiency of datacenters. LEED’s fundamental value proposition is that a novel and re-architected optical network—RotorNet— can deliver “more bandwidth per buck” as well as unique system-level attributes that significantly improve overall datacenter energy efficiency and performance. LEED has developed three system-level testbeds. The first testbed uses calibrated hardware and software power measurements to determine server energy efficiency as a function of network bandwidth and workload. These measurements have shown that increasing network communications bandwidth dramatically increases server energy efficiency providing a realistic path to the overall ENLITENED program goal of doubling the number of transactions per joule. The second testbed demonstrates key hardware: a prototype low-loss, high-port count optical “selector switch”. This switch was fabricated, racked, and tested. Measured switch characteristics include loss, bandwidth, crosstalk, switch time, system-level switch time (including the transceivers), and bit error rate. The third testbed demonstrates a fully working and manufactured pinwheel design which dramatically lowers the cost of design, while delivering high switch radix and low reconfiguration times. The LEED project has tied these three novel photonic switch prototypes together with production servers and software through the development of a novel FPGA-based NIC platform called Corundum. Corundum ensures that the packet-switched protocols supported by commodity operating systems and devices can interface with the Rotor switch design. The LEED group has used this combined hardware and software prototype to characterize applications running at a commercially relevant scale. The project has used a combination of enhanced optical modulation amplitude (OMA) modulators, broadband multiplexers and demultiplexers, avalanche photodiodes, and a novel burst-mode receivers to enable the insertion of LEED-developed optical switches without the need for expensive optical amplification. Our modeling has shown that measured LEED-developed device characteristics can achieve link characteristics of 2 pJ/bit including both transceivers and the Rotor switch. In summary, the LEED program has demonstrated a credible and practical path, through novel hardware and software, to realize the program objectives of ENLITENED. The net result will ensure that the United States maintains its strength in the crucial sector of Information Technology, which is vital to both our economic security and our national security.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Network packet templating for GPU-initiated communication

Systems, apparatuses, and methods for performing network packet templating for graphics processing unit (GPU)-initiated communication are disclosed. A central processing unit (CPU) creates a network packet according to a template and populates a first subset of fields of the network packet with static data. Next, the CPU stores the network packet in a memory. A GPU initiates execution of a kernel and detects a network communication request within the kernel and prior to the kernel completing execution. Responsive to this determination, the GPU populates a second subset of fields of the network packet with runtime data. Then, the GPU generates a notification that the network packet is ready to be processed. A network interface controller (NIC) processes the network packet using data retrieved from the first subset of fields and from the second subset of fields responsive to detecting the notification.

97 MATHEMATICS AND COMPUTING↗