Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Solid State Quantum Refrigeration Superconducting, Absorption and Measurement Based (Final Technical Report)

During this DOE grant, DE-SC0017890, in place for the past six years, all proposed research was carried out and published in peer-reviewed papers, as well as other projects that emerged during the research. In that effort the research team accomplished all proposed research, as well as many closely related research projects discovered and conceived of during the grant. These works included “Efficient Quantum Measurement Engines”, a work published in Physical Review Letters, giving a theory of quantum measurement-based engines, which uses quantum measurement as a resource. These engines are designed to efficiently convert energy from the stochastic quantum measurement process into useful work. Further publications include “Experimental Realization of a Quantum Dot Energy Harvester”, a joint theory and experimental work in collaboration with the group of Charles Smith in Cambridge, UK, as well as long time theoretical collaborators, Rafael Sánchez and Björn Sothmann. This work, featured as an Editor’s Suggestion in Physical Review Letters, realized an earlier theoretical proposal of ours, whereby two resonant tunneling quantum dots are connected to a central electronic cavity that is heated by a hot energy source. We also published “Superconducting Quantum Refrigerator: Breaking and Rejoining Cooper Pairs with Magnetic Field Cycles” a work done in collaboration with experimentalist Francesco Giazotto from ENS Pisa, Italy, which also resulted in a patent. This paper, published in Phys. Rev. Applied, advanced the concept of a cyclic fridge based on the normal/superconducting phase transition together with layered materials separated by tunnel junctions. We also completed the proposed research on a heat transistor, publishing “Thermal transistor and thermometer based on Coulomb-coupled conductors”, carried out as a collaboration between my group and theorists Splettstoesser (Lund U., Sweden), Sothmann (U. Duisburg-Essen, Germany), and Sánchez (U. Autónoma de Madrid, Spain). We carried out an analysis of a quantum coupled to a quantum point contact as a sensitive thermometer and heat transistor. We found the optimal statistical estimator for the temperature and compared it with experiments on the same type of devices. We also investigated autonomous quantum absorption refrigerators using quantum dots to cool by using a very hot thermal reservoir to drive heat between two other reservoirs. In the article “Quantifying the quantum heat contribution from a driven superconducting circuit”, we demonstrated that for a driven superconducting circuit, we showed heat flow provided by a hot source to the qubit can be switched on and off by varying external parameters, the frequency and the intensity of the driving. In the work “Stochastic thermodynamic cycles of a mesoscopic thermoelectric engine”, we reconsidered the autonomous thermoelectric heat engine in terms of underlying cycles. Rather than periodic behavior, the cycles were stochastic in nature. Nevertheless, by undertaking a graph theoretical analysis of the elementary transport processed, great quantitative and qualitative insight could be found. We also considered the quantum measurement process and showed that a quantum version of Maxwell’s demon could be related to the work extraction of a quantum system, closely related to arrow-of-time measures for quantum measurement, as described in our article “Thermodynamics of quantum measurement and Maxwell's demon's arrow of time”. This work was selected in Phys. Rev. A as an Editor’s Suggestion. A recent preprint titled “Cyclic Superconducting Quantum Refrigerators Using Guided Fluxon Propagation” accomplished an important piece of this grant: to propose a new kind of quantum refrigerator using the dynamics of fluxons in a type II superconductor. This invention envisioned a race-track type geometry where fluxons are confined. By applying a gradient of magnetic field together with electric current in a Corbino geometry, the circulating fluxons can actively cool a cold reservoir, realizing a new type of cyclic superconducting refrigerator. We also investigated the possibility of thermal control from different points of view. The application of quantum measurement to the system gives a new kind of control on the system of interest – we have pioneered this approach and shown that measurement can boost the thermal power of quantum engines as described in “Continuous measurement boosted adiabatic quantum thermal machines”. The ability to have heat flows on demand is an outstanding challenge, and we have provided new solutions to this problem in Thermal control across a chain of electronic nanocavities” for a chain of electron cavities using gating voltage control. The control methods using qubit/qubit coupling to create absorption fridges at their most fundamental level have also been developed.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Topology-Informed Design Rules for Deconstructable Thermoset Copolymer Networks

Existing models of thermoset deconstruction facilitated by incorporating cleavable comonomers rely on a mean-field reverse gel point paradigm, which predicts network dissolution once cleavable bonds reach a critical stoichiometric threshold, but does not account for where those bonds reside within the network architecture. Using reactive coarse-grained molecular dynamics simulations coupled with graph-theoretic analysis, we extend this stoichiometric picture to show that deconstructability is governed by the curing-imprinted network topology rather than stoichiometry alone. This topological organization is hierarchical: at the local scale, the elastic effectiveness of cross-link junctions determines which cross-links constitute the load-bearing scaffold; at the mesoscale, the cross-linking rate kinetically templates that scaffold into topologically modular communities─densely cross-linked clusters connected by sparse bridging strands that sustain network connectivity. Using betweenness centrality to identify nodes that disproportionately lie on intercommunity shortest paths, we demonstrate that effective deconstruction of the network into macromolecular fragments requires cleavable comonomers to intercept these high-centrality bridging strands. We further find that under uniform, disassortative comonomer incorporation, this topological requirement provides a mechanistic basis for extending the reverse gel point to incorporate network topology. We also show that modularity imposes a fundamental limit on fragment uniformity that persists even when the centrality requirement is met. Finally, we demonstrate that chain stiffness provides a nearly independent lever to suppress mechanically redundant cross-links and raise the glass transition temperature without significantly altering the deconstruction outcome. Together, these findings reframe the thermoset design space around network topology and provide actionable guidelines for engineering thermoset copolymers with predictable deconstructability and targeted thermomechanical performance.

coarse-grained molecular dynamics↗

Effects of Nonequilibrium Atomic Structure on Ionic Diffusivity in LLZO: A Classical and Machine Learning Molecular Dynamics Study

To improve the performance of electrochemical devices, it is essential to understand the effects of nonequilibrium motifs in solids, such as grain boundaries, amorphous phases, and highly strained regions, on atomic-scale transport and stability. Molecular dynamics simulations are used to explore the combined effect of far-from-equilibrium atomic structures and the choice of interatomic potential on ionic diffusivity predictions for Li 7 La 3 Zr 2 O 12 (LLZO), a promising solid electrolyte for all-solid-state batteries. Amorphization and high strain are considered using both classical Buckingham interatomic potentials and machine learning force fields. Here we find that both crystalline expansion and amorphization tend to slow diffusion, although the different physical encodings in the two potentials impact the properties in different ways. We trace these variations to a combination of structural and transport factors, the contributions of which are deconvoluted computationally. Graph-based analysis reveals that the variations for amorphous LLZO arise from the connectivity of diffusion pathways within the predicted structures, which generally correlates with diffusivity and is notably higher for structures generated by the machine learning force fields. Our study provides additional insight into the relationship between atomic structure and diffusivity in LLZO, while also highlighting the need for care in choosing and validating potentials to simulate far from equilibrium structures.

25 ENERGY STORAGE↗

Dynamic conformational switching underlies TFIIH function in transcription and DNA repair and impacts genetic diseases

Transcription factor IIH (TFIIH) is a protein assembly essential for transcription initiation and nucleotide excision repair (NER). Yet, understanding of the conformational switching underpinning these diverse TFIIH functions remains fragmentary. TFIIH mechanisms critically depend on two translocase subunits, XPB and XPD. To unravel their functions and regulation, we build cryo-EM based TFIIH models in transcription- and NER-competent states. Using simulations and graph-theoretical analysis methods, we reveal TFIIH’s global motions, define TFIIH partitioning into dynamic communities and show how TFIIH reshapes itself and self-regulates depending on functional context. Our study uncovers an internal regulatory mechanism that switches XPB and XPD activities making them mutually exclusive between NER and transcription initiation. By sequentially coordinating the XPB and XPD DNA-unwinding activities, the switch ensures precise DNA incision in NER. Mapping TFIIH disease mutations onto network models reveals clustering into distinct mechanistic classes, affecting translocase functions, protein interactions and interface dynamics.

59 BASIC BIOLOGICAL SCIENCES↗

Minimization of Measurement Uncertainty in Optical Frequency Domain Reflectometry

Optical frequency domain reflectometry (OFDR) is a technique for interrogating optical fiber sensors to generate relative, quasi-distributed measurements. Although Optical frequency domain reflectometry (OFDR) is increasingly being adopted for aerospace, energy production, and structural monitoring applications, the quantification of uncertainty for OFDR measurements has not been developed beyond sparse empirical relationships. To address this knowledge gap, an uncertainty metric for OFDR measurements was developed. This uncertainty metric was applied to weight the edges between OFDR measurements on directed correlation graphs and analyzed to minimize the cumulative uncertainty. In conclusion, this work is the first to propose an uncertainty metric for OFDR and provides a generalized mathematical framework for optimizing OFDR hardware selection, optical fiber sensor selection, and postprocessing strategy.

42 ENGINEERING↗

Structural Controllability Assessment for Inverter-Based Microgrids

Enhanced inverter-based controls are considered for microgrids, which use additional actuation beyond a droop-like term. Shaping of the microgrid’s small-signal dynamics using such enhanced controls is posed as a structural controllability problem. A graph-theoretic characterization of structural controllability is obtained, in terms of the concept of zero-forcing sets.Two benchmark test systems are used to illustrate the selection of locations where enhanced controls should be applied, based on the graph-theoretic analysis. These examples indicate that small-signal characteristics can be shaped using a relatively small number of enhanced controls

microgrids, zero-forcing, structural controlabilit↗

Efficient Topology Assessment for Integrated Transmission and Distribution Network with 10,000+ Inverter-based Resources

The renewable energy proliferation calls upon the grid operators and planners to systematically evaluate the potential impacts of distributed energy resources (DERs). Considering the significant differences between various inverter-based resources (IBRs), especially the different capabilities between grid-forming inverters and grid-following inverters, it is crucial to develop an efficient and effective assessment procedure besides available co-simulation framework with high computation burdens. This paper presents a streamlined graph-based topology assessment for the integrated power system transmission and distribution networks. Graph analyses were performed based on the integrated graph of modified miniWECC grid model and IEEE 8500-node test feeder model, high performance computing platform with 40 nodes and total 2400 CPUs has been utilized to process this integrated graph, which has 100,000+ nodes and 10,000+ IBRs. The node ranking results not only verified the applicability of the proposed method, but also revealed the potential of distributed grid forming (GFM) and grid following (GFL) inverters interacting with the centralized power plants.

Graph Analysis, Topology evaluation, Infrastructur↗

Mechanical coupling in the nitrogenase complex

The enzyme nitrogenase reduces dinitrogen to ammonia utilizing electrons, protons, and energy obtained from the hydrolysis of ATP. Mo-dependent nitrogenase is a symmetric dimer, with each half comprising an ATP-dependent reductase, termed the Fe Protein, and a catalytic protein, known as the MoFe protein, which hosts the electron transfer P-cluster and the active-site metal cofactor (FeMo-co). A series of synchronized events for the electron transfer have been characterized experimentally, in which electron delivery is coupled to nucleotide hydrolysis and regulated by an intricate allosteric network. We report a graph theory analysis of the mechanical coupling in the nitrogenase complex as a key step to understanding the dynamics of allosteric regulation of nitrogen reduction. This analysis shows that regions near the active sites undergo large-scale, large-amplitude correlated motions that enable communications within each half and between the two halves of the complex. Computational predictions of mechanically regions were validated against an analysis of the solution phase dynamics of the nitrogenase complex via hydrogen-deuterium exchange. These regions include the P-loops and the switch regions in the Fe proteins, the loop containing the residue β-188Ser adjacent to the P-cluster in the MoFe protein, and the residues near the protein-protein interface. In particular, it is found that: (i) within each Fe protein, the switch regions I and II are coupled to the [4Fe-4S] cluster; (ii) within each half of the complex, the switch regions I and II are coupled to the loop containing β-188Ser; (iii) between the two halves of the complex, the regions near the nucleotide binding pockets of the two Fe proteins (in particular the P-loops, located over 130 Å apart) are also mechanically coupled. Notably, we found that residues next to the P-cluster (in particular the loop containing β-188Ser) are important for communication between the two halves.

59 BASIC BIOLOGICAL SCIENCES↗

Altered cortical thickness-based structural covariance networks in type 2 diabetes mellitus

Cognitive impairment is a common complication of type 2 diabetes mellitus (T2DM), and early cognitive dysfunction may be associated with abnormal changes in the cerebral cortex. This retrospective study aimed to investigate the cortical thickness-based structural topological network changes in T2DM patients without mild cognitive impairment (MCI). Fifty-six T2DM patients and 59 healthy controls underwent neuropsychological assessments and sagittal 3-dimensional T1-weighted structural magnetic resonance imaging. Then, we combined cortical thickness-based assessments with graph theoretical analysis to explore the abnormalities in structural covariance networks in T2DM patients. Correlation analyses were performed to investigate the relationship between the altered topological parameters and cognitive/clinical variables. T2DM patients exhibited significantly lower clustering coefficient (C) and local efficiency (Elocal) values and showed nodal property disorders in the occipital cortical, inferior temporal, and inferior frontal regions, the precuneus, and the precentral and insular gyri. Moreover, the structural topological network changes in multiple nodes were correlated with the findings of neuropsychological tests in T2DM patients. Thus, while T2DM patients without MCI showed a relatively normal global network, the local topological organization of the structural network was disordered. Moreover, the impaired ventral visual pathway may be involved in the neural mechanism of visual cognitive impairment in T2DM patients. This study enriched the characteristics of gray matter structure changes in early cognitive dysfunction in T2DM patients.

Huang, Yang↗

Evaluation of global teleconnections in CMIP6 climate projections using complex networks

In climatological research, the evaluation of climate models is one of the central research subjects. As an expression of large-scale dynamical processes, global teleconnections play a major role in interannual to decadal climate variability. Their realistic representation is an indispensable requirement for the simulation of climate change, both natural and anthropogenic. Therefore, the evaluation of global teleconnections is of utmost importance when assessing the physical plausibility of climate projections. We present an application of the graph-theoretical analysis tool δ-MAPS, which constructs complex networks on the basis of spatio-temporal gridded data sets, here sea surface temperature and geopotential height at 500 hPa. Complex networks complement more traditional methods in the analysis of climate variability, like the classification of circulation regimes or empirical orthogonal functions, assuming a new non-linear perspective. While doing so, a number of technical tools and metrics, borrowed from different fields of data science, are implemented into the δ-MAPS framework in order to overcome specific challenges posed by our target problem. Those are trend empirical orthogonal functions (EOFs), distance correlation and distance multicorrelation, and the structural similarity index. δ-MAPS is a two-stage algorithm. In the first place, it assembles grid cells with highly coherent temporal evolution into so-called domains. In a second step, the teleconnections between the domains are inferred by means of the non-linear distance correlation. We construct 2 unipartite and 1 bipartite network for 22 historical CMIP6 climate projections and 2 century-long coupled reanalyses (CERA-20C and 20CRv3). Potential non-stationarity is taken into account by the use of moving time windows. The networks derived from projection data are compared to those from reanalyses. Our results indicate that no single climate projection outperforms all others in every aspect of the evaluation. But there are indeed models which tend to perform better/worse in many aspects. Differences in model performance are generally low within the geopotential height unipartite networks but higher in sea surface temperature and most pronounced in the bipartite network representing the interaction between ocean and atmosphere.

58 GEOSCIENCES↗

Antenna Near-Field Probe Station Scanner

A miniaturized antenna system is characterized non-destructively through the use of a scanner that measures its near-field radiated power performance. When taking measurements, the scanner can be moved linearly along the x, y and z axis, as well as rotationally relative to the antenna. The data obtained from the characterization are processed to determine the far-field properties of the system and to optimize the system. Each antenna is excited using a probe station system while a scanning probe scans the space above the antenna to measure the near field signals. Upon completion of the scan, the near-field patterns are transformed into far-field patterns. Along with taking data, this system also allows for extensive graphing and analysis of both the near-field and far-field data. The details of the probe station as well as the procedures for setting up a test, conducting a test, and analyzing the resulting data are also described.

Zaman, Afroz J.↗

A Cloud-Based Global Flood Disaster Community Cyber-Infrastructure: Development and Demonstration

Flood disasters have significant impacts on the development of communities globally. This study describes a public cloud-based flood cyber-infrastructure (CyberFlood) that collects, organizes, visualizes, and manages several global flood databases for authorities and the public in real-time, providing location-based eventful visualization as well as statistical analysis and graphing capabilities. In order to expand and update the existing flood inventory, a crowdsourcing data collection methodology is employed for the public with smartphones or Internet to report new flood events, which is also intended to engage citizen-scientists so that they may become motivated and educated about the latest developments in satellite remote sensing and hydrologic modeling technologies. Our shared vision is to better serve the global water community with comprehensive flood information, aided by the state-of-the- art cloud computing and crowdsourcing technology. The CyberFlood presents an opportunity to eventually modernize the existing paradigm used to collect, manage, analyze, and visualize water-related disasters.

CyberFlood↗

Understanding Machine Learning in Earth Science: A Natural Language Processing Approach

Machine learning (ML) is being increasingly utilized in Earth science research. Benefits of ML include efficiency, reduction of human error, and ability to extract hidden patterns within data. However, the mutual lack of each other’s domain knowledge by ML and Earth science stands as a barrier to timely and effective implementation. Earth science, in particular, faces challenges in generating sample data, compared to those of traditional ML problems such as face recognition or stock predictions, where data is abundant and not lacking in ground truth, which is necessary for labeling. Earth science data are more varying in formats, such as HDF5 and image resolutions, and are not standardized across instruments, even within a given Earth science discipline. Previous studies have been done to outline the specific challenges that Earth science faces with ML, while others have focused on using existing publications to mine information efficiently. Other resources such as Scikit-Learn have developed decision trees for choosing appropriate machine learning algorithms, but application within Earth science subjects becomes much more complex. For the current study, we propose a methodology and tool that aids in implementation of ML in Earth science using natural language processing (NLP). Our work comprises three main parts: (1) analyzing existing publications related to ML and Earth science, using natural language processing: (2) extracting from the publications information on ML models subjects in Earth Science: and (3) visualizing the extracted relationships as a network graph. The resulting network graph should aid the Earth science communities in applying optimal ML algorithms and guiding data preparation through visualization of similar studies. The network graph and analysis of document similarity will be the basis of our next step, which is to develop a decision tree for selecting optimal machine learning methodologies for specified Earth science applications.

Zheng, Laura↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, whole organism, behavior; tabular, imagery). Open Science is the concept that the more people have access to scientifically curated data, the more knowledge will be gained. This led NASA to start the development of GeneLab in 2015. GeneLab houses spaceflight and space-analog multi-omics datasets from plant, rodent, small animal, and microbial experiments. The success and knowledge gained from GeneLab led to a new alliance of NASA “Open Science Data Repositories” (OSDR), which include the Ames Life Sciences Data Archive (ALSDA) and the NASA Biological Institutional Scientific Collection (NBISC). Both are adopting the GeneLab data system, so data are more findable, accessible, interoperable, and reusable (FAIR). OSDR systems provide users the ability to upload, download, search, share, analyze, and visualize. Open Science also needs strong confidence in the data, which is gained through building science communities. With ~400 current members, GeneLab and ALSDA formed Analysis Working Groups (AWGs) to provide feedback on processing pipelines, metadata curation standards (for ‘omics and phenotypic-physiological-behavioral assays), and to collaborate in effectively reusing data. The AWG also led to the development of the Radiation Biology Ontology (RBO), ensuring radiation metadata are efficiently captured, connected, and interoperable. Feedback from the AWG provided design input toward the new single point-of-entry data submission portal for all investigators to submit, curate, and share their research data. Space biological data is now maximally open access, collected-curated with rich metadata, and formatted for interoperability to enable systems biology, meta-analysis, knowledge graphs, machine learning, modeling, and other reuse approaches. With potential for further federation of OSDR for data mining with traditional biological and medical databases (NIH, NCI, EBI, etc.), a new era for space biology has begun to support the knowledge discovery necessary for Lunar and Martian missions.

Ryan T Scott↗

Increasing Discovery and Usability of Earth Science Satellite Data with My NASA Data

For 20 years, the My NASA Data project at NASA Langley Research Center has developed innovative approaches to increase the use of NASA’s satellite data by learners. My NASA Data offers a variety of authentic Earth Science datasets and a data visualization tool, eliminating the need for educators and/or learners to obtain specialized knowledge of GIS data formats and software to access and use authentic Earth Science data. While there is no shortage of available data, as federal government agencies such as NASA house petabytes of freely accessible Earth Science datasets, much of the data are only available for download and visualization in specialized formats and software, limiting their accessibility to educators and learners, especially those in primary and secondary school. Using the Google Earth Engine platform, the My NASA Data team has recently reinvented their data visualization tool, called the Earth System Data Explorer (ESDE). The ESDE gives users the capability to explore over 60 Earth Science satellite datasets in a multitude of formats such as maps, graphs, and data table Its new and improved user interface design was developed based on the preferences of educators, whom the My NASA Data project has over 20 years’ experience working with. Earth Science and GIS Subject Matter Experts (SMEs) structured the data in a professional and scientific manner. During Fiscal Year 2023, the My NASA Data website received over 1 million digital engagements, with over one-third being visitors to the data visualization tool. These metrics highlight the interest in a visualization tool that is simple and free to use with reliable and trusted datasets. The ESDE empowers users to readily relate and analyze NASA Earth Science data within their area of interest. The team used a user-centered design (UCD) framework to receive and incorporate feedback into the application’s design. Core requested features include the ability to create time series graphs, comparative analysis of maps, and download the data as CSV file. Responses indicate that advances in data visualization tools such as the ESDE make authentic Earth Science data more accessible. This presentation will cover how the My NASA Data project develops tools to enhance data discovery and accessibility, as well as how SME and user suggestions are incorporated.

Desiray Wilson↗

Visualizing Organizational Influence on Energy Infrastructure

Energy Infrastructure components depend on an evolving, interdependent business ecosystem exposed to long-term, legal, adversarial tactics. An INL-Naval Postgraduate School partnership was designed to support INL Lab Directed Research and Development, NPS graduate research projects, and joint publications. The Technology, Organization, and Person of interest Graph Extraction, Analysis, and Reporting (TOP GEAR) enumerates networks of organizations and people that own, operate, and maintain regional infrastructure assets. TOP GEAR allows analysts to model current and future state what-if scenarios that include technological and policy mitigations.

29 - ENERGY PLANNING, POLICY AND ECONOMY↗

SEAFORML (Smart Exploration and Analysis For Optimal and Robust Machine Learning)

The poster discusses data analysis of the WAVgraph database and applied machine learning methods for it. The database is a long-term project that seeks to be a comprehensive repository of information on cyber threats and is updated regularly. It was previously unanalyzed and unexplored. The goal was to learn more about it and its contents in order to have a better understanding and enable better use. The data analysis and discovery enabled further exploration through natural language processing, similarity, and clustering methods. The poster shows some of the insights from the analysis and explains the methods used for the machine learning applications.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗