Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Subgroup analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Are Ferroan Anorthosites Direct Products of the Lunar Magma Ocean?

According to Lunar Magma Ocean (LMO) theory, lunar samples that fall into the ferroan anorthosite (FAN) category represent the only samples we have of of the primordial crust of the Moon. Modeling indicates that plagioclase crystallizes after >70% LMO crystallization and formed a flotation crust, depending upon starting composition. The FAN group of highlands materials has been subdivided into mafic-magnesian, mafic-ferroan, anorthositic- sodic, and anorthositic-ferroan, although it is not clear how these subgroups are related. Recent radiogenic isotope work has suggested the range in FAN ages and isotopic systematics are inconsistent with formation of all FANs from the LMO. While an insulating lid could have theoretically extend the life of the LMO to explain the range of the published ages, are the FAN compositions consistent with crystallization from the LMO? As part of a funded Emerging Worlds proposal (NNX15AH76G), we examine this question through analysis of FAN samples. We compare the results with various LMO crystallization models, including those that incorporate the influence of garnet.

Neal, C. R.↗

Heavy-tailed distribution of the number of papers within scientific journals

Scholarly publications represent at least two benefits for the study of the scientific community as a social group. First, they attest to some form of relation between scientists (collaborations, mentoring, heritage, …), useful to determine and analyze social subgroups. Second, most of them are recorded in large databases, easily accessible and including a lot of pertinent information, easing the quantitative and qualitative study of the scientific community. Understanding the underlying dynamics driving the creation of knowledge in general, and of scientific publication in particular, can contribute to maintaining a high level of research, by identifying good and bad practices in science. In this article, we aim to advance this understanding by a statistical analysis of publication within peer-reviewed journals. Namely, we show that the distribution of the number of papers published by an author in a given journal is heavy-tailed, but has a lighter tail than a power law. Interestingly, we demonstrate (both analytically and numerically) that such distributions match the result of a modified preferential attachment process, where, on top of a Barabási-Albert process, we take the finite career span of scientists into account.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Hindlimb Suspension (HLS) in Rodents for the Study of Intracranial Pressure, Molecular and Histologic Changes in the Eye, and CSF Production Regulation and Resorption: A Status Report of Two Studies

This status report corresponds to two studies tied to an animal experiment being executed at the University of California Davis (Charles Fuller's laboratory). The animal protocol uses the well-documented rat hindlimb suspension (HLS) model, to examine the relationship between cephalic fluid shifts and the regulation of intracranial (ICP) and intraocular (IOP) pressures as well as visual system structure and function. Long Evans rats are subjected to HLS durations of 7, 14, 28 and 90 days. Subgroups of the 90-day animals are studied for recovery periods of 7, 14, 28 or 90 days. All HLS subjects have age-matched cage controls. Various animal cohorts are planned for this study: young males, young females and old males. In addition to the live measures (ICP by telemetry, IOP and retinal parameters by optical coherence tomography) which are shared with the Fuller study, the specific outcomes for this study include: -Gene expression analysis of the retina -Histologic analysis - Analysis of the microvasculature of retina flat mounts by NASA's VESsel GENeration Analysis (VESGEN) Software. To date, the young male and female cohorts are being completed. Due to the need to keep technical variation to a minimum, the histologic and genomic analyses have been delayed until all samples from each cohort are available and can be processed in a single batch per cohort. The samples received so far correspond to young males sacrificed at 7,14, 28 and 90 days of HLS and at 90 days of recovery; and from young females sacrificed at 7, 14 and 28 of HLS. A complementary study titled: "A gene expression and histologic approach to the study of cerebrospinal fluid (CSF) production and outflow in hindlimb suspended rats" seeks to study the molecular components of CSF production and outflow modulation as a result of HLS, bringing a molecular and histologic approach to investigate genome wide expression changes in the arachnoid villi and choroid plexus of HLS rats compared to rats in normal posture.

Theriot, C. A.↗

Classification of mafic clasts from mesosiderites - Implications for endogenous igneous processes

Results are presented from an analysis of 13 igneous pebbles from the Vaca Muerta, EET87500, and Bondoc mesosiderites, using electron microprobe and instrumental neutron activation techniques. These data, combined with literature data on compositions of 43 mesosiderite clasts were used to compile a classification scheme for the various types of mafic silicate clasts that occur in mesosiderites. These clasts were classified into five principal groups: (1) polygenic and monogenic cumulates (30 percent); (2) polygenic basalts (30 percent); (3) quench-textured rocks, comprising two compositional subgroups (those which resemble basaltic eucrites (5 percent), and those which resemble cumulate eucrites (2 percent)); (4) monogenic basalts (11 percent); and (5) ultramafic rocks, consisting mainly of large crystals of orthopyroxene (9 percent) or olivine (4 percent). The conditions under which these clasts were formed are discussed.

Rubin, Alan E.↗

A detailed petrological analysis of hydrated, low-nickel, nonchondritic stratospheric dust particles

A detailed petrological analysis of three low-Ni, K-bearing, nonchondritic stratospheric dust particles is performed, and these particles are compared to products of high-energy, explosive (Plinian-type) volcanic events. The analytical electron microscope (AEM) analyses show pervasive layer silicates, carbonate and goethite, and chemical fractionation in the matrix of these particles similar to hydrothermal alteration in volcanic ejecta. Along with low Ni content and the presence of potassium, the texture and mineralogy of particles L2001-18, L2001-20, and L2002 C2 are similar to at least two nonchondritic stratospheric dust particles of the igneous subgroup for which an extraterrestrial origin has been suggested based on their minor- and trace-element abundances. The petrological characteristics of some low-Ni, K-bearing nonchondritic stratospheric dust particles supports a probable terrestrial volcanic origin, but the AEM data alone cannot exclude an extraterrestrial origin for these particles.

Rietmeijer, Frans J. M.↗

Archaebacterial phylogeny: perspectives on the urkingdoms

Comparisons of complete 16S ribosomal RNA sequences have been used to confirm, refine and extend earlier concepts of archaebacterial phylogeny. The archaebacteria fall naturally into two major branches or divisions, I--the sulfur-dependent thermophilic archaebacteria, and II--the methanogenic archaebacteria and their relatives. Division I comprises a relatively closely related and phenotypically homogeneous collection of thermophilic sulfur-dependent species--encompassing the genera Sulfolobus, Thermoproteus, Pyrodictium and Desulfurococcus. The organisms of Division II, however, form a less compact grouping phylogenetically, and are also more diverse in phenotype. All three of the (major) methanogen groups are found in Division II, as are the extreme halophiles and two types of thermoacidophiles, Thermoplasma acidophilum and Thermococcus celer. This last species branches sufficiently deeply in the Division II line that it might be considered to represent a separate, third Division. However, both the extreme halophiles and Tp. acidophilum branch within the cluster of methanogens. The extreme halophiles are specifically related to the Methanomicrobiales, to the exclusion of both the Methanococcales and the Methanobacteriales. Tp. acidophilum is peripherally related to the halophile-Methanomicrobiales group. By 16S rRNA sequence measure the archaebacteria constitute a phylogenetically coherent grouping (clade), which excludes both the eubacteria and the eukaryotes--a conclusion that is supported by other sequence evidence as well. Alternative proposals for archaebacterial phylogeny, not based upon sequence evidence, are discussed and evaluated. In particular, proposals to rename (reclassify) various subgroups of the archaebacteria as new kingdoms are found wanting, for both their lack of proper experimental support and the taxonomic confusion they introduce.

NASA Discipline Exobiology↗

Measuring Equality in Machine Learning Security Defenses: A Case Study in Speech Recognition

Over the past decade, the machine learning security community has developed a myriad of defenses for evasion attacks. An understudied question in that community is: for whom do these defenses defend? This work considers common approaches to defending learned systems and how security defenses result in performance inequities across different sub-populations. We outline appropriate parity metrics for analysis and begin to answer this question through empirical results of the fairness implications of machine learning security methods. We find that many methods that have been proposed can cause direct harm, like false rejection and unequal benefits from robustness training. The framework we propose for measuring defense equality can be applied to robustly trained models, preprocessing-based defenses, and rejection methods. We identify a set of datasets with a user-centered application and a reasonable computational cost suitable for case studies in measuring the equality of defenses. In our case study of speech command recognition, we show how such adversarial training and augmentation have non-equal but complex protections for social subgroups across gender, accent, and age in relation to user coverage. We present a comparison of equality between two rejection-based defenses: randomized smoothing and neural rejection, finding randomized smoothing more equitable due to the sampling mechanism for minority groups. This represents the first work examining the disparity in the adversarial robustness in the speech domain and the fairness evaluation of rejection-based defenses.

• Artificial intelligence (AI) / machine learning ↗

Correlations of Prompt and Afterglow Emission in Swift Long and Short Gamma Ray Bursts

Correlation studies of prompt and afterglow emissions from gamma-ray bursts (GRBs) between different spectral bands has been difficult to do in the past because few bursts had comprehensive and intercomparable afterglow measurements. In this paper we present a large and uniform data set for correlation analysis based on bursts detected by the Swift mission. For the first time, short and long bursts can be analyzed and compared. It is found for both classes that the optical, X-ray and gamma-ray emissions are linearly correlated, but with a large spread about the correlation line; stronger bursts tend to have brighter afterglows, and bursts with brighter X-ray afterglow tend to have brighter optical afterglow. Short bursts are, on average, weaker in both prompt and afterglow emissions. No short bursts are seen with extremely low optical to X-ray ratio as occurs for 'dark' long bursts. Although statistics are still poor for short bursts, there is no evidence yet for a subgroup of short bursts with high extinction as there is for long bursts. Long bursts are detected in the dark category at the same fraction as for pre-Swift bursts. Interesting cases are discovered of long bursts that are detected in the optical, and yet have low enough optical to X-ray ratio to be classified as dark. For the prompt emission, short and long bursts have different average tracks on flux vs fluence plots. In Swift, GRB detections tend to be fluence limited for short bursts and flux limited for long events.

Gehrel, Neil↗

Streptococcus pneumoniae HtrA is a dynamic and monomeric virulence factor capable of forming larger oligomeric complexes

Abstract High‐temperature requirement A (HtrA) proteases are a conserved family of serine proteases central to protein quality control and bacterial virulence. While Gram‐negative and human HtrAs are structurally well studied, Gram‐positive homologs remain essentially uncharacterized. Here, we present the first integrated structural and mechanistic analysis of a Gram‐positive HtrA, from Streptococcus pneumoniae , a virulence factor essential for adhesion and infection in vivo. Proteomic profiling of an htrA knockout and cleavage assays demonstrate that S. pneumoniae HtrA is required for protein quality control, with the PDZ domain mediating substrate recognition. Biochemically, S. pneumoniae HtrA exists exclusively as a monomer in solution, a striking divergence from canonical trimeric HtrAs that we show is shared with other Gram‐positive homologs. NMR analyses reveal that the monomer dynamically samples open and closed conformations, while cryo‐EM of a catalytic mutant identifies a hexamer stabilized by a unique LoopA–PDZ interaction. Together, these findings define S. pneumoniae HtrA as a dynamic monomer with interdomain coupling between its protease and PDZ domains, establishing Gram‐positive HtrAs as a mechanistically divergent subgroup within the HtrA family.

Lee, Eunjeong [Department of Biochemistry and Mole↗

Chimeric calcium/calmodulin-dependent protein kinase in tobacco: differential regulation by calmodulin isoforms

cDNA clones of chimeric Ca2+/calmodulin-dependent protein kinase (CCaMK) from tobacco (TCCaMK-1 and TCCaMK-2) were isolated and characterized. The polypeptides encoded by TCCaMK-1 and TCCaMK-2 have 15 different amino acid substitutions, yet they both contain a total of 517 amino acids. Northern analysis revealed that CCaMK is expressed in a stage-specific manner during anther development. Messenger RNA was detected when tobacco bud sizes were between 0.5 cm and 1.0 cm. The appearance of mRNA coincided with meiosis and became undetectable at later stages of anther development. The reverse polymerase chain reaction (RT-PCR) amplification assay using isoform-specific primers showed that both of the CCaMK mRNAs were expressed in anther with similar expression patterns. The CCaMK protein expressed in Escherichia coli showed Ca2+-dependent autophosphorylation and Ca2+/calmodulin-dependent substrate phosphorylation. Calmodulin isoforms (PCM1 and PCM6) had differential effects on the regulation of autophosphorylation and substrate phosphorylation of tobacco CCaMK, but not lily CCaMK. The evolutionary tree of plant serine/threonine protein kinases revealed that calmodulin-dependent kinases form one subgroup that is distinctly different from Ca2+-dependent protein kinases (CDPKs) and other serine/threonine kinases in plants.

NASA Discipline Plant Biology↗

A Model of Reduced Kinetics for Alkane Oxidation Using Constituents and Species for N-Heptane

The reduction of elementary or skeletal oxidation kinetics to a subgroup of tractable reactions for inclusion in turbulent combustion codes has been the subject of numerous studies. The skeletal mechanism is obtained from the elementary mechanism by removing from it reactions that are considered negligible for the intent of the specific study considered. As of now, there are many chemical reduction methodologies. A methodology for deriving a reduced kinetic mechanism for alkane oxidation is described and applied to n-heptane. The model is based on partitioning the species of the skeletal kinetic mechanism into lights, defined as those having a carbon number smaller than 3, and heavies, which are the complement of the species ensemble. For modeling purposes, the heavy species are mathematically decomposed into constituents, which are similar but not identical to groups in the group additivity theory. From analysis of the LLNL (Lawrence Livermore National Laboratory) skeletal mechanism in conjunction with CHEMKIN II, it is shown that a similarity variable can be formed such that the appropriately non-dimensionalized global constituent molar density exhibits a self-similar behavior over a very wide range of equivalence ratios, initial pressures and initial temperatures that is of interest for predicting n-heptane oxidation. Furthermore, the oxygen and water molar densities are shown to display a quasi-linear behavior with respect to the similarity variable. The light species ensemble is partitioned into quasi-steady and unsteady species. The reduced model is based on concepts consistent with those of Large Eddy Simulation (LES) in which functional forms are used to replace the small scales eliminated through filtering of the governing equations; in LES, these small scales are unimportant as far as the overwhelming part of dynamic energy is concerned. Here, the scales thought unimportant for recovering the thermodynamic energy are removed. The concept is tested by using tabular information from the LLNL skeletal mechanism in conjunction with CHEMKIN II utilized as surrogate ideal functions replacing the necessary functional forms. The test reveals that the similarity concept is indeed justified and that the combustion temperature is well predicted, but that the ignition time is over-predicted, a fact traced to neglecting a detailed description of the processes leading to the heavies chemical decomposition. To palliate this deficiency, functional modeling is incorporated into this conceptual reduction in addition to the modeling the evolution of the global constituent molar density, the enthalpy evolution of the heavies, the contribution to the reaction rate of the unsteady lights from other light species and from the heavies, the molar density evolution of oxygen and water, and the mole fractions of the quasisteady light species. The model is compact in that there are only nine species-related progress variables. Results are presented showing the performance of the model for predicting the temperature and species evolution. The model reproduces the ignition time over a wide range of equivalence ratios, initial pressure, and initial temperature.

Harstad, Kenneth G.↗

Domain Shift Analysis in Chest Radiographs Classification in a Veterans Healthcare Administration Population

This study aims to assess the impact of domain shift on chest X-ray classification accuracy and to analyze the influence of ground truth label quality and demographic factors such as age group, sex, and study year. We used a DenseNet121 model pre-trained MIMIC-CXR dataset for deep learning-based multi-label classification using ground truth labels from radiology reports extracted using the CheXpert and CheXbert Labeler. We compared the performance of the 14 chest X-ray labels on the MIMIC-CXR and Veterans Healthcare Administration chest X-ray dataset (VA-CXR). The validation of ground truth and the assessment of multi-label classification performance across various NLP extraction tools revealed that the VA-CXR dataset exhibited lower disagreement rates than the MIMIC-CXR datasets. Additionally, there were notable differences in AUC scores between models utilizing CheXpert and CheXbert. When evaluating multi-label classification performance across different datasets, minimal domain shift was observed in the unseen VA dataset, except for the label “Enlarged Cardiomediastinum.” The subgroup with the most significant variations in multi-label classification performance was study year. These findings underscore the importance of considering domain shift in chest X-ray classification tasks, paying particular attention to the temporality of the exam. Our study reveals the significant impact of domain shift and demographic factors on chest X-ray classification, emphasizing the need for improved transfer learning and robust model development. Addressing these challenges is crucial for advancing medical imaging research and improving patient care.

chest X-ray image classification↗

GAL08, an Uncultivated Group of Acidobacteria, Is a Dominant Bacterial Clade in a Neutral Hot Spring

GAL08 are bacteria belonging to an uncultivated phylogenetic cluster within the phylum Acidobacteria . We detected a natural population of the GAL08 clade in sediment from a pH-neutral hot spring located in British Columbia, Canada. To shed light on the abundance and genomic potential of this clade, we collected and analyzed hot spring sediment samples over a temperature range of 24.2–79.8°C. Illumina sequencing of 16S rRNA gene amplicons and qPCR using a primer set developed specifically to detect the GAL08 16S rRNA gene revealed that absolute and relative abundances of GAL08 peaked at 65°C along three temperature gradients. Analysis of sediment collected over multiple years and locations revealed that the GAL08 group was consistently a dominant clade, comprising up to 29.2% of the microbial community based on relative read abundance and up to 4.7 × 10 5 16S rRNA gene copy numbers per gram of sediment based on qPCR. Using a medium quality threshold, 25 single amplified genomes (SAGs) representing these bacteria were generated from samples taken at 65 and 77°C, and seven metagenome-assembled genomes (MAGs) were reconstructed from samples collected at 45–77°C. Based on average nucleotide identity (ANI), these SAGs and MAGs represented three separate species, with an estimated average genome size of 3.17 Mb and GC content of 62.8%. Phylogenetic trees constructed from 16S rRNA gene sequences and a set of 56 concatenated phylogenetic marker genes both placed the three GAL08 bacteria as a distinct subgroup of the phylum Acidobacteria , representing a candidate order ( Ca. Frugalibacteriales) within the class Blastocatellia. Metabolic reconstructions from genome data predicted a heterotrophic metabolism, with potential capability for aerobic respiration, as well as incomplete denitrification and fermentation. In laboratory cultivation efforts, GAL08 counts based on qPCR declined rapidly under atmospheric levels of oxygen but increased slightly at 1% (v/v) O 2 , suggesting a microaerophilic lifestyle.

59 BASIC BIOLOGICAL SCIENCES↗

SPARC: Structural properties associated with residue constraints

SPARC facilitates the generation of plausible hypotheses regarding underlying biochemical mechanisms by structurally characterizing protein sequence constraints. Such constraints appear as residues co-conserved in functionally related subgroups, as subtle pairwise correlations (i.e., direct couplings), and as correlations among these sequence features or with structural features. SPARC performs three types of analyses. First, based on pairwise sequence correlations, it estimates the biological relevance of alternative conformations and of homomeric contacts, as illustrated here for death domains. Second, it estimates the statistical significance of the correspondence between directly coupled residue pairs and interactions at heterodimeric interfaces. Third, given molecular dynamics simulated structures, it characterizes interactions among constrained residues or between such residues and ligands that: (a) are stably maintained during the simulation; (b) undergo correlated formation and/or disruption of interactions with other constrained residues; or (c) switch between alternative interactions. We illustrate this for two homohexameric complexes: the bacterial enhancer binding protein (bEBP) NtrC1, which activates transcription by remodeling RNA polymerase (RNAP) containing σ 54 , and for DnaB helicase, which opens DNA at the bacterial replication fork. Based on the NtrC1 analysis, we hypothesize possible mechanisms for inhibiting ATP hydrolysis until ADP is released from an adjacent subunit and for coupling ATP hydrolysis to restructuring of σ 54 binding loops. Based on the DnaB analysis, we hypothesize that DnaB ‘grabs’ ssDNA by flipping every fourth base and inserting it into cavities between subunits and that flipping of a DnaB-specific glutamine residue triggers ATP hydrolysis.

97 MATHEMATICS AND COMPUTING↗

Two new extremely hot pulsating white dwarfs

High speed photometry of the extremely hot, nearly degenerate stars PG 1707 + 427 and PG 2131 + 066 reveals that they are low-amplitude pulsating variables. Power spectral analysis shows both to be multiperiodic, with dominant periods of 7.5 and 6.4-6.9 minutes, respectively. Together with the known pulsators PG 1159 - 035 and the central star of the planetary nebula Kohoutek 1-16, these objects define a new pulsational instability strip at the hot edge of the H-R diagram. The variations of these objects closely resemble those of the much cooler pulsating ZZ Ceti DA white dwarfs; both groups are probably nonradial g-mode pulsators. Evolutionary contraction of the PG 1159 - 035 variables may lead to period changes that would be detectable in as little as 1 year. The optical and IUE spectra of the PG 1159 - 035 variables are characterized by absorption lines of C IV and other CNO ions, indicating radiative levitation of species heavier than helium. He II is also present in the spectra, but the hydrogen Balmer lines are absent. Effective temperatures near 100,000 K are required, and the He II 4686 A profiles indicate log g greater than 6. These helium-rich pulsators form the hottest known subgroup of the DO white dwarfs.

Bond, H. E.↗

SN 2021fxy: mid-ultraviolet flux suppression is a common feature of Type Ia supernovae

ABSTRACT We present ultraviolet (UV) to near-infrared (NIR) observations and analysis of the nearby Type Ia supernova SN 2021fxy. Our observations include UV photometry from Swift/UVOT, UV spectroscopy from HST/STIS, and high-cadence optical photometry with the Swope 1-m telescope capturing intranight rises during the early light curve. Early B − V colours show SN 2021fxy is the first ‘shallow-silicon’ (SS) SN Ia to follow a red-to-blue evolution, compared to other SS objects which show blue colours from the earliest observations. Comparisons to other spectroscopically normal SNe Ia with HST UV spectra reveal SN 2021fxy is one of several SNe Ia with flux suppression in the mid-UV. These SNe also show blueshifted mid-UV spectral features and strong high-velocity Ca ii features. One possible origin of this mid-UV suppression is the increased effective opacity in the UV due to increased line blanketing from high velocity material, but differences in the explosion mechanism cannot be ruled out. Among SNe Ia with mid-UV suppression, SNe 2021fxy and 2017erp show substantial similarities in their optical properties despite belonging to different Branch subgroups, and UV flux differences of the same order as those found between SNe 2011fe and 2011by. Differential comparisons to multiple sets of synthetic SN Ia UV spectra reveal this UV flux difference likely originates from a luminosity difference between SNe 2021fxy and 2017erp, and not differing progenitor metallicities as suggested for SNe 2011by and 2011fe. These comparisons illustrate the complicated nature of UV spectral formation, and the need for more UV spectra to determine the physical source of SNe Ia UV diversity.

Astronomy & Astrophysics↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗

Substructure in the stellar halo near the Sun: I. Data-driven clustering in integrals-of-motion space

Context. Merger debris is expected to populate the stellar haloes of galaxies. In the case of the Milky Way, this debris should be apparent as clumps in a space defined by the orbital integrals of motion of the stars. Aims. Our aim is to develop a data-driven and statistics-based method for finding these clumps in integrals-of-motion space for nearby halo stars and to evaluate their significance robustly. Methods. We used data from Gaia EDR3, extended with radial velocities from ground-based spectroscopic surveys, to construct a sample of halo stars within 2.5 kpc from the Sun. We applied a hierarchical clustering method that makes exhaustive use of the single linkage algorithm in three-dimensional space defined by the commonly used integrals of motion energy E, together with two components of the angular momentum, L z and L ⊥ . To evaluate the statistical significance of the clusters, we compared the density within an ellipsoidal region centred on the cluster to that of random sets with similar global dynamical properties. By selecting the signal at the location of their maximum statistical significance in the hierarchical tree, we extracted a set of significant unique clusters. By describing these clusters with ellipsoids, we estimated the proximity of a star to the cluster centre using the Mahalanobis distance. Additionally, we applied the HDBSCAN clustering algorithm in velocity space to each cluster to extract subgroups representing debris with different orbital phases. Results. Our procedure identifies 67 highly significant clusters (> 3σ), containing 12% of the sources in our halo set, and 232 subgroups or individual streams in velocity space. In total, 13.8% of the stars in our data set can be confidently associated with a significant cluster based on their Mahalanobis distance. Inspection of the hierarchical tree describing our data set reveals a complex web of relations between the significant clusters, suggesting that they can be tentatively grouped into at least six main large structures, many of which can be associated with previously identified halo substructures, and a number of independent substructures. This preliminary conclusion is further explored in a companion paper, in which we also characterise the substructures in terms of their stellar populations. Conclusions. Our method allows us to systematically detect kinematic substructures in the Galactic stellar halo with a data-driven and interpretable algorithm. The list of the clusters and the associated star catalogue are provided in two tables available at the CDS.

79 ASTRONOMY AND ASTROPHYSICS↗