MainMuch of the protein universe remains unexplored, with little known about sequence or structural similarities to studied proteins. As illustrated in Fig. 1a, only 571,864 (0.2%) protein sequences have been curated and reviewed manually by UniProtKB/Swiss-Prot7, leaving the vast majority—more than 245 million—annotated computationally and assigned functional or family membership on the basis of sequence homology. As depicted in Fig. 1b, this concept of darkness, or the ‘dark proteome’, also extends to human proteins, with 5,516 out of 20,412 (27%) classified as the darkest category, Tdark (poorly characterized with minimal functional data, which lack known approved drugs), Tclin (drug-targeted), Tchem (potent small molecule ligands, for example, high-potency small molecule binders) and Tbio (well-studied biology (disease relevant but of unclear druggability))8.Fig. 1: Identifying superdark GPCR-like proteins.a, Most proteins in UniProt (99.5%) remain understudied (only 0.2% are reviewed), with functions inferred mainly from sequence homology or computational predictions, representing the ‘dark matter’ of the protein universe. b, Darkness in the human proteome is categorized into four levels: Tclin, Tchem, Tbio and Tdark. c, 7TMPs were identified through an exhaustive structural similarity search using the GPCR rhodopsin structure (Protein Data Bank (PDB) 1F88) as a query. UniProt and InterPro classifications were used to distinguish AF2 matches as GPCRs or 7TMPs by superkingdom: A (archaeal), B (bacterial), E (eukaryotic) or U (unclassified). d, AF2-predicted 7TMP structures aligned to the rhodopsin template (white) trimmed to display only their 7TM domains, highlighting the human repertoire of sensor, sensor-like and non-sensor 7TMPs and superdark candidates. Models are coloured by predicted local distance difference test (pLDDT) score from blue (low) to red (high). e,f, Tree-like network graphs spanning 2,000 representative 7TMPs, coloured by node depth from parent (e) and superkingdom (f). The corresponding interactive Cytoscape networks are provided as Supplementary Data 2 and 3, respectively. HK, histidine kinase; HRH1, histamine H1 receptor; SMO, Smoothened. g, Successive node-to-node sequence conservation across 7TM architectures originating in archaea and culminating in eukaryotic GPCRs. UniProt identifiers are provided in d,g. K, thousand; M, million.Source dataIn this work, we expand the concept of the dark human proteome to include a category of understudied proteins beyond Tdark. We define these proteins as superdarks and hypothesize that they share structural homology with proteins of known function while lacking definitive sequence similarity to well-characterized proteins. By leveraging advanced structure prediction tools such as Alphafold4,5, we can now identify superdarks with untapped therapeutic potential and deepen our understanding of cell and molecular biology in unexpected ways. Here we show how this approach uncovered an ancient GPCR-like superdark protein with these characteristics, underscoring the promise of exploring the understudied human proteome and beyond.Superdark 7TM families resemble GPCRsSequence-based homology searches primarily use BLAST9 and hidden Markov model-based approaches such as HMMER10 to identify evolutionary relationships. However, these approaches are constrained by their reliance on detectable sequence similarity, often called the ‘twilight zone’1,2,3, making it challenging to recognize deeply sequence-divergent proteins with similar structures and functions. By contrast, structure-based approaches such as DALI11, CE12, TM-align13,14 and Foldseek15 exploit the principle that structure is more conserved than sequence over evolutionary timescales16 to identify remote homologues, even when sequence similarity is minimal or undetectable.Seven transmembrane proteins (7TMPs) fulfill various roles across all domains of life. They function as GPCRs in eukaryotic signalling, adhesion molecules in cell–cell interactions, ion pumps in microbial energy metabolism, chemoreceptors in bacterial chemotaxis and environmental sensors in stress responses. Recognizing their central role in biology, we conducted an exhaustive pairwise structure-based homology search for superdark 7TMPs against the 214,528,851 predictions in the Alphafold2 (AF2) database. As shown in Fig. 1c and Extended Data Fig. 3, we used TM-align and the prototypical GPCR rhodopsin as a query structure to identify 1,543,898 initial matches.Because matching overall three-dimensional (3D) shape is only the first step in identifying structural homologues, we next developed structure-based heuristics adapted from our previously reported pHinder algorithms17,18,19,20 to rank the quality of each match. Rather than removing poor matches, these geometric filters distinguished full-length, topologically correct folds from partial or low-quality alignments and prioritized the most complete structural homologues. After removing 82,419 AF2 models with obsolete UniProt identifiers, we assigned each of the remaining 1,461,479 hits a rank of 1, 2 or 3. Rank 1 proteins exhibited full fold coverage, rank 2 proteins contained N-terminal or C-terminal truncations or internal coverage gaps, and rank 3 proteins displayed spurious or low structural coverage (Extended Data Fig. 3a). Overall, 47%, 36% and 17% of hits fell into ranks 1, 2 and 3, respectively (Supplementary Data 1).We next classified the ranked proteins as GPCRs or non-GPCR 7TMPs using the UniProtKB and InterPro21 databases. As shown in Fig. 1c, 35% (506,129) of AF2 predictions correspond to known GPCRs spanning 15,368 organisms (Supplementary Data 3). The remaining 65% (955,350) represent non-GPCR 7TMPs distributed across 371,568 organisms (Supplementary Data 4), with robust representation across all three superkingdoms. Together, these results provide an unprecedented collection of GPCR and non-GPCR 7TMP structural homologues. As illustrated in Fig. 1d, representative examples include canonical sensors such as GPCRs, which transduce signals through G proteins, histidine kinases, which detect changes in pH, osmolarity and chemical signals, and diguanylate cyclases (DGCs), which sense nutrient availability and stress signals to regulate bacterial biofilm formation, as well as structurally related sensor-like and non-sensor proteins. Among these are representative human superdark GPCR-like candidates, including members of the TM184 and PRRT families.The structural relationships revealed by our search prompted us to ask whether these proteins also retain detectable sequence similarity. Conceptually, this represents an ‘inverse AlphaFold’ problem. Rather than asking whether sequence predicts structure, we asked whether proteins sharing a common seven transmembrane (7TM) fold retain sufficient sequence similarity to reveal their relationship. To address this, we compiled sequences for all 1,461,479 7TMPs and applied our geometric trimming procedure to isolate the 7TM core. Redundancy was reduced with MMseqs2 (ref. 22) to obtain 39,808 representative sequences, which were compared in an all-versus-all BLASTp to construct a sequence similarity network. We then traversed this network using a breadth-first search initiated from the highest-degree node, enforcing thresholds of at least 30% identity and a bitscore of 30 or greater, which we found to be the minimum values required to recover known GPCR homologies. Benchmarking against 2,029 human 7TMP and GPCR AF2 models (Supplementary Data 1) showed that stricter thresholds would exclude known GPCRs, validating our cut-off choice.Inspection of these networks revealed that initiating from the highest-degree archaeal node (a histidine kinase) produced tree-like structures spanning several 7TMP families (Fig. 1e). In the superkingdom-coloured view (Fig. 1f), GPCR-like clusters comprised proteins from archaea, bacteria and eukaryotes, underscoring the taxonomic breadth of the 7TM fold. Strikingly, traversal from archaeal histidine kinases to eukaryotic GPCRs required no more than six sequential similarity steps, showing that only a few degrees of separation connect archaeal 7TMPs to GPCR families (Fig. 1g). Together, these analyses show that GPCR-like folds can be encoded by many primary sequences that share only minimal sequence signals.More than 800 human GPCRs regulate key physiological processes and are principal drug targets, yet more than 100 remain classified as Tdark. We next asked whether structural similarity could reveal an even darker class of GPCR-like proteins. Our search identified eight candidates, independent of the GPCR query seed, whereas Foldseek results varied depending on the seed structure used (Extended Data Fig. 4). We excluded TMEM116 and TMEM187 because they have been proposed to encode GPCRs21. The remaining six candidates belong to two 7TMP families: OSTα/TM184 (TM184A, TM184B, TM184C and SLC51A) and PRRT (PRRT3 and PRRT4). Representative members, TM184C and PRRT3, are highlighted in Fig. 1d.To prioritize candidates for experimental characterization, we first evaluated the expression of PRRT and OSTα/TM184 family members across human tissues and cell types using the Human Protein Atlas23,24. As shown in Extended Data Fig. 1a, the expression pattern of PRRT4 remains unknown, whereas PRRT3, SLC51A and TM184A exhibit modest expression across diverse tissues and cell types. By contrast, TM184B and TM184C are expressed more broadly, with TM184C displaying the highest and most ubiquitous expression among the superdark candidates. Combined with its Tdark classification, these findings led us to focus our subsequent studies on TM184C.We next characterized TM184C expression and subcellular localization in HEK293T cells, comparing it with the other superdark candidates (TM184A/B, SLC51A and PRRT3/4), the Golgi-localized 7TMP TMEM181 and representative GPCRs (AVPR2 and ADRB2) (Extended Data Fig. 1b–d). To ensure plasma membrane trafficking of SLC51A, we co-expressed its obligate partner SLC51B25, and split luciferase complementation detected low but measurable surface expression (Extended Data Fig. 1c). Bystander bioluminescence resonance energy transfer (BRET) measurements across seven subcellular compartments corroborated these findings and further revealed that TM184C, despite robust expression, exhibited minimal plasma membrane localization (Extended Data Fig. 1c, d), resembling endoplasmic reticulum-retained SLC51A and Golgi-localized TMEM181. By contrast, TM184A/B and AVPR2 were detected readily at the cell surface, although substantial fractions of TM184A/B also localized to recycling and early endosomes, suggesting a role in endocytic trafficking. TM184C, by comparison, localized predominantly to late endosomes and lysosomes (Extended Data Fig. 1d).Superdarks recruit arrestin and GRKWe next investigated whether TM184C contained structural features associated with GPCR activation and G protein coupling. As expected, it lacked class A (LxxD, CWxP, PIF, NPxxY, E\DRY) and class B (HETx, PxxG) motifs26,27,28. However, TM184C harboured an NRY sequence (N135–R136–Y137) at a position corresponding to the E\DRY motif in TM3 of many class A GPCRs (Extended Data Fig. 1e), including our query rhodopsin structure (E134–R135–Y136). Despite modest sequence similarity to other OSTα/TM184 family members (Extended Data Figs. 1e and 5), this NRY motif was unique to TM184C. Although the NRY motif of TM184C aligned with the ERY of rhodopsin, the presence of glutamine in TM6 prevented an ionic lock, leaving the evolutionary relevance of its NRY sequence unclear.As shown in Extended Data Figs. 1e and 5a, we next searched for GPCR-like features beyond the 7TM core. We focused on arrestin codes6, which are patterns of serine (S) and threonine (T) residues in GPCR C termini that are phosphorylated by GPCR kinases (GRKs) to regulate arrestin recruitment. Except for SLC51A, all candidate superdarks contained these codes. Within the OSTα/TM184 family, TM184C had four arrestin codes (Extended Data Fig. 1e), whereas TM184A and B had one and two, respectively (Extended Data Fig. 5a). By contrast, PRRT3/4 displayed six arrestin codes (Extended Data Fig. 5a). These observations suggested that the superdark candidates TM184C and PRRT3/4 may recruit arrestin.Assuming that evolutionary age reflects sustained biological importance, we next searched for the oldest TM184C and PRRT3/4 proteins using PSI-BLAST9. Our analysis revealed that, whereas PRRT proteins are confined to eukaryotes, TM184C can be traced back to archaea and bacteria through direct homology (Extended Data Figs. 1f,g and 5d). To the best of our knowledge, GPCRs cannot be traced to archaea or bacteria by sequence alone, underscoring the uniqueness of TM184C and suggesting a potential evolutionary link to primordial GPCR-like folds. Moreover, the predicted canonical GPCR-like orthosteric site, combined with a non-canonical intracellular effector site (Extended Data Figs. 1h and 5e), calculated using pHinder17,18,19,20 and Fpocket29, reinforced the idea that TM184C may have structural hallmarks of primordial GPCR architectures.As shown in Fig. 2, we next tested TM184C for GPCR features using BRET, pulldown and immunoblotting assays to assess its recruitment of G proteins, arrestins and GRKs. Without a known ligand, we relied on TM184C constitutive activity and used the 4A assay, which detects active-state G-protein coupling in the absence of ligand stimulation30. In Fig. 2a, none of the OSTα/TM184 or PRRT family proteins, including TM184C, recruited active-state 4AG proteins, whereas controls such as MRGPRD and ADRB2 produced robust BRET signals. We obtained similar results using mini G31 and TRUPATH32 assays (Fig. 2b–d). Although TM184C exhibited a slightly stronger BRET loss than our Gs (ADRB2) and Gi (DRD2) controls, these results were modest and not reproducibly above background. On the basis of these observations, we conclude that TM184C does not show clear evidence of constitutive G protein coupling under our experimental conditions, although weak or context-dependent interactions cannot be excluded.Fig. 2: TM184C associates with canonical GPCR effectors.a, 4A assay with RLuc8-tagged superdarks co-expressed with mVenus-tagged Gβγ subunits, where BRET increase indicates Gα recruitment. b, Recruitment of mVenus-tagged mini Gα subunits to RLuc8-tagged receptors, with BRET increase indicating G-protein recruitment; data are normalized to Hras. c, TRUPATH assay using RLuc8-tagged Gα and GFP2-tagged βγ subunits, where BRET decrease indicates G-protein dissociation. d, Constitutive TM184C dose-dependent Gs dissociation in the TRUPATH assay compared with a Gs (ADRB2) non-Gs (DRD2) coupling control. e, Arrestin recruitment measured by co-expressing RLuc8-tagged receptors with C-terminally mVenus-tagged arrestins, where a gain of BRET indicates binding. f, Comparison of arrestin recruitment between TM184C wild-type and Δ82 mutant (left to right: P = 0.2096, 0.3564, 0.0021, 0.4738 and 0.0600). g, Co-immunoprecipitation (Co-IP) of FLAG–TM184C (WT and Δ82) with ARRB1–3xHA and ARRB2–3xHA in HEK293T cells, immunoblotted (IB) for FLAG and HA. h, GRK co-expression increases arrestin recruitment to TM184C. i, GRK2–mVenus recruitment to RLuc8-tagged receptors measured by BRET. j, Phosphorylation sites on TM184C (pink) identified by mass spectrometry of the indicated SDS–PAGE band align with predicted arrestin motifs (blue). k, Co-IP of FLAG–TM184C WT and ACM with ARRB1–3xHA. l, GRK-dependent phosphorylation of TM184C confirmed by λ-phosphatase (λ-phos) treatment. βarr, β-arrestin. P values were calculated using an unpaired, two-tailed Welch’s t-test. Data in a–d, e, f, h and i are the mean; d, f and i also show s.d. (error bars), n = 4 independent experiments for a–d and n = 3 for e, f, h and i. **P ≤ 0.01; NS, not significant.Source dataHaving shown that TM184C is unlikely to recruit G proteins, we next examined its ability to recruit other canonical GPCR effectors, starting with arrestins. Using BRET assays with mVenus-tagged wild-type arrestins (ARR1 and ARR2) and activated arrestin variants (ARRB1-FL and ARRB2-K78E), we found that TM184C was the only OSTα/TM184 family member to recruit arrestins at levels comparable with basal ADRB2 arrestin binding (Fig. 2e). We observed similar arrestin recruitment in the PRRT family, confirming the identification of two new families of GPCR-like proteins, although we maintained our focus on TM184C.Arrestins can engage GPCRs through the receptor core or the C-terminal tail. To test whether arrestins interact with the C terminus of TM184C, we generated a TM184C mutant lacking its C-terminal domain (TM184C-Δ82) and observed reduced ARRB2 recruitment (Fig. 2f). We next performed immunoprecipitation experiments (Fig. 2g): more full-length TM184C co-immunoprecipitated with both ARRB1 and ARRB2, whereas TM184C-Δ82 pulled down less arrestin, reinforcing the view that the C-terminal tail of TM184C is the primary site mediating arrestin recruitment. Finally, given that GRKs often enhance arrestin binding by phosphorylating GPCR C termini, we repeated our BRET experiments for full-length TM184C using our arrestin panel co-expressed with individual GRKs (2, 3, 5 and 6) (Fig. 2h). We observed that all GRKs increased arrestin recruitment to TM184C; this was most pronounced with GRK2 and when using wild-type arrestins. These results suggested that TM184C is phosphorylated in a GRK-dependent manner and that phosphorylation promotes stronger arrestin binding.Having established a connection between GRKs and TM184C arrestin recruitment, we focused on GRK2, which showed the most pronounced effect (Fig. 2h). Using BRET, we measured GRK2 recruitment to TM184C and several control receptors (Fig. 2i). Full-length TM184C recruits GRK2 as robustly as ADRB2. By contrast, the truncated variant TM184C-Δ82 did not recruit GRK2, consistent with our arrestin recruitment data (Fig. 2f,g), confirming that the sites critical for GRK and arrestin interactions reside in the C terminus of TM184C. We next identified the phosphorylation sites in its C terminus using proteomics. We detected three sites—S422, S432 and S435—all of which mapped to predicted arrestin codes (Fig. 2j and Extended Data Fig. 1e). We then mutated these and eight other phosphorylation sites in the four arrestin codes (S374, S377, S379, S380, S382, S424, S427 and S438) to alanine to create the TM184C-arrestin code mutant (ACM) variant. As shown in Fig. 2k, TM184C-ACM pulled down less ARRB1, demonstrating that the phosphorylation sites within the arrestin codes provide the molecular mechanism for enhanced arrestin recruitment.As shown in Fig. 2k, TM184C migrated as a doublet when arrestin codes were present and as a singlet when they were absent, suggesting that the upper band represents phosphorylated TM184C. To test this, we examined TM184C phosphorylation across our full panel of GRKs in the presence and absence of λ-phosphatase (Fig. 2l). GRK2 and GRK6 produced the most pronounced phosphorylation, and λ-phosphatase eliminated the upper TM184C band in every case, confirming its phosphorylation-dependent mobility shift. Together, these results demonstrate that GRKs phosphorylate the TM184C C-terminal arrestin codes. Thus, TM184C exhibits two of the three canonical GPCR properties: β-arrestin recruitment and GRK-dependent phosphorylation of its C-terminal arrestin-binding sites. These signalling features resemble atypical chemokine receptors, which signal primarily through arrestins rather than through G proteins33,34. Accordingly, TM184C and related proteins may represent a previously unrecognized class of arrestin-coupled receptors, rather than canonical GPCRs.TM184C marks intracellular vesiclesHaving characterized several facets of TM184C biochemistry, we next assessed its subcellular localization, extending our previous bystander BRET experiments (Extended Data Fig. 1d) using confocal microscopy to visualize endogenous TM184C using immunofluorescence and live-cell imaging of a TM184C–eGFP fusion expressed stably in HEK293A cells. As shown in Fig. 3, TM184C is not detected on the plasma membrane, contrary to expectations for a canonical GPCR. Instead, TM184C localizes to vesicular structures ranging in size from approximately 500 nm to 5 μm. These vesicles sample the nuclear envelope dynamically, are embedded within the nucleus and surveil the cytoplasm (Fig. 3a,b) and are frequently enriched in perinuclear clusters and narrow cellular projections (Fig. 3b).Fig. 3: TM184C localizes to intracellular vesicles and promotes the formation of intercellular connections.a, Representative confocal images of endogenous TM184C detected by immunofluorescence using a TM184C-specific antibody show vesicular distribution within the cell; nuclei are stained with 4′,6-diamidino-2-phenylindole (DAPI; blue) and TM184C vesicles are magenta. b, Live-cell imaging of HEK293A cells expressing TM184C tagged with eGFP reveals a similar vesicle distribution. c, Confocal live-cell imaging demonstrates TM184C–eGFP vesicles (magenta) localized along microtubules stained with Tubulin Tracker Deep Red (cyan), with sequential frames (0–18 s) showing bidirectional movement, including vesicular scission and redistribution. d,e, TM184C vesicle localization within cellular projections is illustrated in cells expressing TM184C–eGFP, showing vesicles protruding into two short projections (approximately 5 µm and 23 µm in length) (d), and depicting vesicles travelling and aggregating at the ends of longer projections (extending from about 50 µm to 58 µm within 3 s) (e), with insets displaying distal region. f, Time-lapse imaging captures TM184C vesicle dynamics during the formation of intercellular connections between two cells, with sequential frames (0–28 s) showing vesicle movement along microtubules until a connection is established. g–i, TM184C vesicles and acidic vesicle transport within cellular compartments and intercellular connections; nuclei are stained in blue (NucBlue), TM184C–eGFP in magenta and acidic vesicles with LysoTracker Red are green. Distribution of TM184C vesicles concentrating in the perinuclear region, intercellular connections and projections (g); LysoTracker-stained acidic vesicles with a similar distribution pattern but in lower numbers (h) and overlay with magnified views revealing TM184C vesicles enveloping acidic vesicles (i), suggesting spatial and functional interaction. Representative examples are shown for n = 3 independent experiments. See companion Supplementary Videos 1–5. Scale bars, 1 μm (insets in a–c; top two insets in g–i), 5 μm (a–d; insets in e, bottom insets in g–i), 2 μm (insets in f), 10 μm (e–i).The unexpected vesicular localization is well aligned with our GRK- and β-arrestin-related findings (Fig. 2e–l). The C-terminal tail of TM184C contains arrestin code motifs6 and undergoes phosphorylation consistent with GRK-mediated modification and subsequent β-arrestin recruitment—a canonical mechanism that drives internalization, placing TM184C within the endolysosomal system rather than at the cell surface. To identify which specific vesicular compartments contain TM184C, we examined its co-localization with a panel of mVenus-tagged subcellular markers (Extended Data Fig. 1d). As shown in Extended Data Fig. 6, late endosome and lysosome markers co-localized with TM184C–eGFP vesicles, reaffirming our BRET findings. We also observed robust co-localization of LC3B–mCherry35 with TM184C–eGFP, demonstrating that TM184C also resides in autophagosomes (Extended Data Fig. 6). Consistent with these assignments, LysoTracker staining revealed that many, but not all, TM184C–eGFP vesicles are acidic. These findings establish that TM184C partitions into both acidified and non-acidified vesicle pools within the late endocytic and autophagic pathways.Having established TM184C localization to late endosomal and autophagic vesicles, we next examined their dynamics using live-cell imaging. As shown in Fig. 3c and Supplementary Video 1, TM184C vesicles move rapidly along microtubules, undergoing frequent fission and fusion events during transit. These vesicles exhibit robust bidirectional movement along individual microtubules, consistent with coordinated engagement of kinesin motors for anterograde transport and dynein motors for retrograde transport. Immunofluorescence further demonstrates strong co-localization of TM184C vesicles with KIF5B (Extended Data Fig. 6k), a kinesin family member that mediates organelle positioning and long-range transport of mitochondria and lysosomes. Building on these dynamics, we noted that TM184C vesicles are highly enriched in thin, microtubule-supported projections in cell lines stably overexpressing TM184C. As shown in Fig. 3d, e, Extended Data Fig. 7b and Supplementary Videos 2 and 3, we observed these projections and their underlying microtubule networks, along which TM184C vesicles are distributed and extend to distal reaches of the cell, suggesting that TM184C overexpression enhances projection formation and maintenance—a relationship that we substantiate further in rescue experiments presented below.TM184C drives intercellular transferHaving established the association of TM184C with cellular projections, we next observed that TM184C-positive extensions can fuse to form intercellular bridges (Fig. 3f–i). These structures resemble tunnelling nanotubes36,37,38 and tumour microtubes39, which serve as conduits for long-range cell–cell communication. As shown in Fig. 3f and Supplementary Videos 4 and 5, TM184C-positive vesicles accumulated at projection tips and nascent junctions where neighbouring projections met, suggesting a role in establishing new intercellular contacts and multicellular networks. In cells with established intercellular bridges, TM184C-positive vesicles were enriched in the perinuclear region and distributed along the length of the connections (Fig. 3g). A subset of TM184C-positive vesicles was LysoTracker-positive, and these vesicles frequently associated closely with, or enveloped, acidic vesicles within both the cell body and intercellular bridges (Fig. 3i), revealing a close spatial relationship between TM184C-positive vesicles and acidic organelles in established intercellular connections. This organization suggests that TM184C-positive bridges support the intercellular transfer of lysosomes and other organelles between connected cells.We next asked whether cells require TM184C to form projections and make intercellular connections. CRISPR-mediated knockdown of TM184C in HEK293A and T cells caused severe morphological and lysosomal defects, as cells became rounded, lost characteristic projections and intercellular bridges, displayed enlarged and mislocalized acidic vesicles and showed abnormal vesicular organization (Extended Data Fig. 7a). TM184C-deficient cells were fragile, poorly adherent and hypersensitive to imaging, often blebbing and dying almost immediately upon exposure to the microscope laser (Extended Data Fig. 7a and Supplementary Video 6), consistent with impaired cytoskeletal and endolysosomal function. By contrast, cells overexpressing TM184C–eGFP or TM184C–mCherry developed extensive projections and interconnections spanning tens to hundreds of micrometres (Extended Data Fig. 7b).To determine how TM184C abundance and its C terminus regulate cellular connectivity, we performed subtractive rescue experiments. Because CRISPR-mediated deletion of TM184C produced fragile cells (Fig. 4a and Extended Data Fig. 7), precluding conventional knockout-and-rescue approaches, we generated HEK293A lines stably expressing TM184C–eGFP or the C-terminal variants TM184C-ACM–eGFP and TM184C-Δ82–eGFP using TM184C coding sequences recoded to prevent recognition by TM184C-targeting guide RNAs (gRNAs). This strategy enabled selective deletion of endogenous TM184C while preserving exogenous expression, allowing us to isolate the contributions of TM184C abundance and its β-arrestin/GRK-interacting C-terminal motifs to bridge and projection formation (Fig. 4a–d). To quantify these phenotypes, we developed an image-analysis pipeline that segments cell bodies and clusters, traces cellular boundaries and interfaces, and identifies bridges, projections and protrusions (Fig. 4a–d and Extended Data Fig. 8). These measurements were integrated into a bridge–projection–protrusion (bpp) score that weights each structure by its maximal extension from the nearest cell body and the number of cells or cell clusters it connects.Fig. 4: TM184C overexpression drives persistent bridge and projection connectivity even after endogenous TM184C knockout.a–d, Representative confocal images of HEK293A cells expressing empty vector (EV) (a), stably overexpressed (OE) TM184C–eGFP (b), TM184C-ACM–eGFP (ACM) (c) or TM184C-Δ82–eGFP (Δ82, C-terminal truncation) (d) with or without TM184C-targeting gRNA1 or gRNA2. Nuclei (NucBlue), blue; LysoTracker, green; and, in the OE, ACM and Δ82 lines, the stably expressed TM184C construct, magenta. In the bridges, projections and protrusions identification overlays, cell or cell-cluster outlines used to calculate the bpp score are traced in yellow, and bridges, projections and protrusions extending from these outlines are pseudo-coloured in orange. See Extended Data Fig. 8 for details of the bpp score calculation. TM184C OE produces a striking bridge/projection-rich phenotype, whereas WT, ACM and Δ82 cells show few bridges and projections when endogenous TM184C is removed. e, bpp scores at baseline (EV) confirm that TM184C OE markedly increases bridges and projections compared with HEK293A controls (one-way analysis of variance (ANOVA); NS comparisons, left to right, P = 0.3295 and 0.4951). f, bpp scores across all genotypes and CRISPR conditions show that TM184C OE maintains high connectivity even when endogenous TM184C is knocked out, whereas WT, ACM and Δ82 cells lose bridges and connections (two-way ANOVA); the ‘rescue’ bracket marks this maintained connectivity. g, Within-genotype comparisons (EV, gRNA1, gRNA2) show that only TM184C OE preserves the high-connectivity phenotype after knockout, whereas ACM and Δ82 fail to rescue bridges and projections (one-way ANOVA). Lines, mean; error bars, s.d.; n = 28–60 regions of interest (ROIs) per condition, pooled from two independent biological replicates, with outliers removed by ROUT (Q = 1%); exact n in the source data (e–g). Statistical comparisons are shown as NS, not significant; ****P < 0.0001. Scale bars, 30 μm.Source dataAs shown in Fig. 4b, cells stably overexpressing TM184C–eGFP formed the most abundant and longest bridges and projections (Fig. 4e–g), with bpp scores approximately threefold above baseline (Fig. 4b versus 4a). This highly connected phenotype persisted after knockout of endogenous TM184C, with bpp scores remaining approximately 15-fold higher than in baseline knockouts. By contrast, HEK293A empty vector controls (Fig. 4a) and HEK293A cells stably expressing TM184C-ACM–eGFP or TM184C-Δ82–eGFP (Fig. 4c,d) exhibited fewer and shorter projections and connections, resulting in markedly lower bpp scores at baseline and near-complete loss of these structures upon CRISPR-mediated deletion of endogenous TM184C (Fig. 4e–g). These findings indicate that TM184C overexpression is sufficient to drive elevated bpp scores and a highly connected morphology, and that intact GRK- and arrestin-recognition motifs within the full TM184C C terminus are required for TM184C-dependent formation of intercellular connections.We next asked whether the intercellular connections driven by TM184C facilitate vesicle transfer between cells. We generated companion HEK293A lines in which TM184 family members or TM184C variants carried C-terminal fusions to either eGFP or mCherry, and paired these lines such that genuine intercellular transfer would be revealed by eGFP-positive vesicles appearing in mCherry-expressing cells and vice versa. To quantify intercellular mixing, we developed a custom triangulation pipeline that detects and visualizes transferred TM184C vesicles directly from confocal fluorescence images (Fig. 5a–e and Extended Data Fig. 9). Vesicles in each colour channel are thresholded, reduced to centroid point clouds and used to construct a triangulation of neighbouring vesicles in co-culture, and short mesh edges are classified as same-colour or mixed-colour connections, yielding an objective, automated measure of vesicles that have moved from one cell into another and of vesicles that now carry both fluorophores.Fig. 5: TM184C and its C terminus facilitate intercellular vesicle transfer.a–e, HEK293A co-cultures stably expressing TM184C–mCherry and TM184C–eGFP (a), TM184C–mCherry and TM184A–eGFP (b), TM184C–mCherry and TM184B–eGFP (c), TM184C–mCherry and TM184C-ACM–eGFP (d) and TM184C–mCherry and TM184C-Δ82–eGFP (e) mixed in pairwise combinations and imaged to assess vesicle transfer. For each pairing, representative ROIs are shown as the overlay (NucBlue-stained nuclei plus both fluorophores) (column 1), individual TM184–mCherry and TM184–eGFP vesicle channels (columns 2 and 3), the pruned triangulation with all nearest-neighbour vesicle edges (same-colour plus mixed) (column 4) and the subset of mixed edges only (column 5), which report TM184 vesicle exchange between cells. The automated triangulation pipeline (Extended Data Fig. 9) detects both pure TM184–mCherry or TM184–eGFP donor vesicles that have entered a recipient cell and vesicles containing both fluorophores, consistent with mixing or fusion of TM184-positive vesicle pools. f, Schematic summary of vesicle flow for each co-culture condition in a–e. TM184C ↔ TM184C pairs exhibit robust bidirectional transfer, whereas TM184C → TM184A, TM184C → TM184B and TM184C → TM184C-ACM combinations show predominantly unidirectional delivery from TM184C-expressing cells. Co-cultures with TM184C-Δ82 display apparent bidirectional exchange, but Δ82 is mislocalized to the plasma membrane (Extended Data Fig. 10), endocytosed and routed into endolysosomal compartments, so its ‘transfer’ reflects degradative passenger vesicles rather than active Δ82-mediated trafficking. The asterisk marks this apparent, non-productive exchange (Supplementary Videos 7–9, 18–21 and 23). Scale bars, 15 μm (d) and 30 μm (a–c,e).As shown in Fig. 5a and Extended Data Fig. 10, co-cultures of TM184C–mCherry and TM184C–eGFP exhibited robust bidirectional vesicle exchange, with each cell line acting as both donor and recipient. By contrast, pairing TM184C–mCherry with TM184A–eGFP or TM184B–eGFP (Fig. 5b, c and Extended Data Fig. 10e–g), which reside primarily at the plasma membrane rather than in vesicles, produced predominantly unidirectional transfer from TM184C-expressing donor cells into TM184A- or TM184B-expressing recipient cells. We next tested how the TM184C C terminus determines this transfer behaviour. When we co-cultured TM184C–mCherry with TM184C-ACM–eGFP (Fig. 5d), vesicle exchange again flowed largely from TM184C donors into ACM-expressing recipients, consistent with a requirement for intact arrestin code motifs and their phosphorylation (Fig. 5d and Extended Data Fig. 10j–k). ACM-expressing cells frequently exhibited altered morphology with enlarged and disorganized vesicles, but co-culture with overexpressed TM184C cells restored vesicle organization and connectivity in ACM cells, indicating that incoming TM184C-positive vesicles can partially rescue these defects (Extended Data Fig. 10k).As shown in Fig. 5e, TM184C–mCherry and TM184C-Δ82–eGFP co-cultures likewise produced mixed-colour vesicles. However, the Δ82 mutant, which lacks the C terminus, mislocalized to the plasma membrane, was internalized and routed into the endolysosomal system for degradation (Fig. 5e and Extended Data Fig. 10h–i). The Δ82-positive structures that transfer to TM184C–eGFP cells therefore remain in degradative compartments, indicating that apparent Δ82 ‘transfer’ reflects mixing of shared lysosomes rather than active trafficking (Extended Data Fig. 10h–i). Together, these experiments establish that TM184C directs the directionality and fate of intercellular vesicle trafficking through its C-terminal arrestin code, as summarized in Fig. 5f.To demonstrate the generality of TM184C-mediated transfer, we extended our co-culture experiments to additional cell types (Extended Data Fig. 10a–d). TM184C–eGFP and TM184C–mCherry donor–recipient pairs in non-cancerous cells (HFF-1 fibroblasts and astrocytes), mixed non-cancerous and cancerous settings (astrocytes with patient-derived glioblastoma (GBM)), and cancer cells (SK-MEL-28 melanoma) all formed TM184C-positive connections and exchanged vesicles (Extended Data Fig. 10a–d and Supplementary Videos 9–13). Finally, we asked whether TM184C vesicles interact with other organelles, focusing on mitochondria because they are also known to transfer between cells40,41,42. TM184C-positive vesicles seemed to closely track and monitor mitochondria in HEK293A cells, astrocytes and patient-derived GBM cells, remaining tightly associated with them both in intercellular connections and within the cell body, and frequently engaged in kiss-and-go contacts (Extended Data Fig. 10l–n and Supplementary Videos 14–16). These observations raise the possibility that TM184C-positive vesicles mediate or assist transcellular metabolic support36,43. In summary, vesicle transfer emerges as a specific property of TM184C, not TM184A or TM184B, and requires an intact, phosphorylated TM184C C terminus, consistent with a model in which β-arrestin- and GRK-dependent GPCR-like activities regulate TM184C trafficking and function to establish multicellular networks.Ancient TM184C constrains autophagyHaving established that TM184C mediates the sharing of vesicular and organellar cargo between cells, we next asked how its localization to late endosomes, lysosomes and autophagosomes might influence cellular recycling pathways. We focused on autophagosomes because TM184C not only localizes to these structures but also contains several C-terminal LC3B-binding motifs44, consistent with a potential link to LC3B, a central component and widely conserved marker of autophagy. LC3B exists in two forms: LC3B-I, a cytosolic precursor, and LC3B-II, a lipidated form that anchors to autophagosomal membranes.Guided by this rationale, we revisited our CRISPR knockdown experiments and analysed LC3B by immunoblotting. The cytosolic (LC3B-I) and lipidated (LC3B-II) forms were resolved readily, revealing that TM184C knockdown in HEK293T cells led to a marked accumulation of LC3B-II (Extended Data Fig. 2a), indicative of increased autophagic activity. This was further supported by the accumulation of abnormally large acidic vesicles upon TM184C knockdown (Extended Data Fig. 7a). Together, these findings suggest that TM184C functions as a negative regulator of autophagy.To determine whether TM184C regulates autophagosome formation or degradation, we next treated knockdown cells with concanamycin A (conA)—a potent vacuolar H+-ATPase inhibitor that prevents lysosomal acidification and blocks autolysosome-mediated cargo degradation. If LC3B-II accumulation reflected defective degradation, conA would cause little or no additional increase45; conversely, if TM184C loss enhanced autophagosome formation, LC3B-II levels would rise further. As shown in Extended Data Fig. 2b, LC3B-II increased markedly following conA treatment, indicating that autophagic flux remained active and that, primarily, TM184C knockdown enhances autophagosome biogenesis rather than impairing clearance. Consistently, both conA and the related H+-ATPase inhibitor bafilomycin A1 (BAF) caused a striking accumulation of enlarged TM184C- and LC3B-positive vesicles (Extended Data Fig. 2d and Supplementary Video 24).Having established that TM184C regulates autophagy, we next asked whether this regulatory role is conserved evolutionarily. Our sequence analyses traced the TM184 family to its archaeal origins (Extended Data Figs. 1f,g and 5d), suggesting that its function pre-dates the emergence of multicellular life. To test this, we examined the yeast homologue Hfl1, which has been implicated previously in vacuolar organization and autophagy. As shown in Extended Data Fig. 2c,e, deleting HFL1 in Saccharomyces cerevisiae reproduced the characteristic multivesicular vacuole phenotype indicative of disrupted autophagic body homeostasis46. Remarkably, CRISPR knock-in of human TM184C fully rescued vacuolar morphology, whereas TM184 A and TM184B did not. Because yeast vacuoles are functional analogues of mammalian lysosomes, and autophagic bodies correspond to autolysosomes, this rescue demonstrates that TM184C retains a deeply conserved role in autophagic regulation. Consistent with this, Hfl1 interacts with Atg8 (ref. 47), the yeast homologue of LC3B. Together, these findings reveal that the structure and function of TM184C are ancient and conserved, linking archaeal membrane systems to modern eukaryotic pathways governing vesicle trafficking and autophagy.DiscussionThis study illustrates how large-scale, structure-first discovery can expose deeply conserved biology hidden beyond sequence similarity. By mining millions of AF2 models with geometry-based filtering, we uncovered TM184C as an ancient GPCR-like regulator of vesicle dynamics, autophagy and intercellular communication. Although our analysis relied on predicted rather than experimentally determined structures, and some candidates may be missed where structural predictions are less reliable, benchmarking against known GPCRs and extensive functional validation in human and yeast systems support the fidelity of our assignments. Future structural studies, particularly cryo electron microscopy or crystallography of TM184C and its complexes, will help define its conformational states and clarify how phosphorylation reshapes its effector interfaces.Functionally, TM184C lies at the intersection of vesicle trafficking, cellular recycling and intercellular exchange. It localizes to late endosomes, lysosomes and autophagosomes, where it constrains autophagic flux. In parallel, TM184C-positive vesicles move along microtubules, concentrate in projections and assemble intercellular bridges through which vesicles, lysosomes, autophagosomes and mitochondria are exchanged. Phosphorylation of its C-terminal arrestin codes by GRKs appears to tune these behaviours and link canonical GPCR-like β-arrestin and GRK machinery to the formation and remodelling of vesicle networks. The ability of human TM184C to restore autophagic body homeostasis in yeast lacking Hfl1 shows that this regulatory logic is both ancient and conserved.Together, these properties suggest that TM184C is a central organizer of multicellular networks of exchange that rely on tunnelling nanotube-like and tumour microtube-like connections—a process for which few dedicated regulators have been identified. Given that β-arrestin is known to scaffold both cofilin/LIMK for actin remodelling at membrane protrusions and kinesin motors for directed cargo transport48,49, the dependence of the TM184C C-terminal tail on intercellular connections and transfer suggests a working model in which β-arrestin recruitment to TM184C coordinates projection formation, microtubule-dependent vesicle transport and autophagosomal membrane dynamics.Viewed more broadly, this study illustrates how integrative approaches that couple structure prediction with quantitative cell biology and evolutionary analysis can illuminate previously hidden aspects of protein function within the dark proteome. The strategy we outline here begins with structural mining for superdark candidates and then links them systematically to cellular phenotypes. This general framework should be broadly applicable for uncovering diverse classes of understudied proteins, including those that regulate processes very different from vesicle trafficking, autophagy or multicellular networks, thereby deepening our understanding of cellular biology and helping to identify new therapeutic targets.MethodsDistributed computingMost calculations in this study were performed on the Triton and Pegasus computing clusters at the Frost Institute for Data Science and Computing at the University of Miami. An overview of our computational pipeline is shown in Extended Data Fig. 3a, indicating the calculation phases that required distributed computing. Except for TM-alig13, we developed the algorithms for this study in our laboratory using Python, integrating structural informatics approaches from our previous work17,18,19,20,50. We submitted our compute jobs using the IBM load-sharing facility (LSF) platform and bsub commands wrapped in custom Python scripts.Downloading the AF2 v.4 datasetOn 20 July 2023, we downloaded the AlphaFold database5 to our high-performance computing environment using the protocol provided by the Google-Deepmind GitHub page51. The download took approximately 1 week, and the compressed archive files (.tar) required 24 tebibytes (TiB) of disk space.Preparing the AF2 database for distributed computingThe 24 TiB AF2 database contained 1,015,798 single-species archive files (.tar) of various sizes. Using a Python script, we redistributed these files into 607 equally sized 40-gibibyte (GiB) groupings of compressed mmCIF files (.cif.gz). We decompressed the mmCIF files in each 40-GiB grouping using a second Python script and distributed computing. Each 40-gigabyte (GB) grouping contained an average of 353,425 mmCIF files (.cif), occupying an average of 100-GiB on disk. The 607 uncompressed AF2 prediction files occupied collectively 60 TiB on disk. Using a Python script, we next equally distributed the 214,528,851 AF predictions of the 607 groupings into 1,000 job submission files.Identifying superdark GPCR-like foldsAs shown in Extended Data Fig. 3a, we performed an exhaustive search for 7TM folds in the 214,528,851 AF2 predictions using the prototypical GPCR rhodopsin (PDB 1F88, chain A)52 as our query structure and the program TM-align13. Using the 1,000 job submission files described above, we generated 1,000 tabular output files (TM-align option ‘-outfmt 2’), each containing pairwise alignment results for 214,528,851 AF predictions. We processed this set of 1,000 output files using a Python script, identifying 1,543,898 matches with at least 200 residues and a TM-align score of 0.5. Next, we identified and removed programmatically 82,419 matches corresponding to the obsolete UniProt database7 entries as of 1 August 2024, arriving at 1,461,479 AF2 structure matches. Because a TM-align score alone is not a robust metric for matching fold topologies, we had to develop additional structure-based algorithms to rank structural similarity and discern the correct 3D fold topology (Extended Data Fig. 3a).Aligning hitsTo generate and visualize structure-based alignments of related protein models, we developed a two-part Python workflow that automated large-scale pairwise superpositions using TM-align. The first script distributed alignment tasks across the Pegasus computing cluster by dividing the input FASTA file into evenly sized subsets (6,000 sequences per job), creating input directories, and submitting array jobs to the cluster scheduler. Each job invoked a second script, which parsed TM-align output files to extract TM-scores, root mean square deviation values and file paths for each query–target pair. For every pair, the script re-ran TM-align to generate the corresponding 3 × 3 rotation and translation matrices, applied these transformations to the atomic coordinates of the mobile structure and wrote the resulting aligned coordinates to new PDB files labelled by alignment score. This automated system enabled efficient, high-throughput structural superposition and normalization of AlphaFold-predicted models for subsequent comparative analysis and visualization.Surface-constrained geometric trimming and topology scoringAs outlined in Extended Data Fig. 3a, for each query–hit pair, structures were parsed, preserving residue indices, chain identifiers and atomic coordinates. We then computed a geometric surface representation of the query, yielding a topology-based spatial mask. Using this mask, we cropped the hit model to the subset of residues whose coordinates fall within, or coincide with, the query surface, thereby defining trimming strictly by geometric overlap rather than by sequence span. The trimmed model and the query surface were compared subsequently with quantify coverage over the query surface and summarize topological discrepancies (for example, internal gaps and N-/C-terminal truncations within the overlapped region). These surface-anchored calculations provided the inputs for downstream filtering and ranking of structural hits.We assign each hit to one of three confidence ranks (r1 > r2 > r3) with a hierarchical decision tree that prioritizes intact, contiguous, high-coverage overlap of the query surface. For every hit we assemble a key comprising coverage, TM-score, hit file and query file, then evaluate three quality gates in order: terminal truncation, coverage gaps and overall coverage. First, if either the N or C terminus is truncated by at least 15% of the overlapped region, the hit cannot be r1. When coverage is at least 0.5 in such cases, the hit is demoted to r2 with a reason flag (–NTC for N-terminal, –CTC for C-terminal); if coverage is less than 0.5 it is classified as r3 with –LC (low coverage). Second, if there are coverage gaps within the overlapped region, we examine their cumulative extent. A cumulative gap of more than 50% forces r3 (–LC), and 15–50% yields r2 (–ICG, incomplete coverage gaps). If gaps are present but small (≤15%), a hit with coverage greater than 0.7 may still qualify as r1 (this is the ‘least stringent’ r1 path under the gaps branch); otherwise it falls to r2 (–ICG). Third, independent of truncation/gap checks, coverage less than 0.7 precludes r1: 0.5–0.7 maps to r2 (–LC) and less than 0.5 maps to r3 (–LC). Hits that clear all three gates—no substantial terminal truncation, no or minimal gaps and coverage of at least 0.7—are assigned r1 by the ‘most stringent’ fall-through path.Constructing the annotated 7TMP catalogue of hitsAs outlined in Extended Data Fig. 3b, we built a master catalogue of sequences and annotations with a four-step Python pipeline, using flat database files for UniProt and InterPro downloaded on 1 September 2025 (offline parsing; no web application programming interfaces). First, raw listings of 7TMP hits were parsed to standardize identifiers (UniProt accession and entry name), capture sequence length and source file paths, normalize heterogeneous headers and remove duplicates, yielding a clean listing table keyed by accession. Second, UniProt flat files (1 September 2025 download) were parsed to extract accession numbers, entry/protein/gene names, organism (and/or TaxID), length and review status, and to harmonize canonical versus isoform records into a metadata table. Third, the listing and UniProt tables were key-merged on accession, preserving provenance paths and adding basic quality control fields (for example, length deltas, missing-metadata flags) to produce a consolidated master table. Finally, InterPro domain annotations from flat files (1 September 2025 download) were integrated by aggregating all protein-to-InterPro mappings per accession and joining them into the master table. The result is a single, analysis-ready CSV/TSV file that unifies cleaned listings, UniProt metadata and InterPro domain sets for downstream filtering, stratification and structural analyses.GPCR-focused filtering and labelling of 7TMP hitsAs outlined in Extended Data Fig. 3b, starting from the annotated catalogue, we applied a two-step rule-based filter to classify 7TMP hits by GPCR status. For each UniProt accession, the script consolidated InterPro signatures and UniProt metadata, then (1) identified GPCR-positive entries using curated GPCR evidence terms and (2) assigned all remaining entries—those not meeting (1)—to a non-GPCR 7TMP set. Results are written as delimited tables and, when enabled, also exported to Excel using a streaming writer suitable for large outputs. This step yields a curated GPCR versus 7TMP partition for downstream analyses.Surface-trimmed sequence extractionTo convert full-length proteins into sequences that reflect only the geometry retained after surface trimming, we generated gap-aware sequences from the trimmed structural models. For each model, Cα records were traversed in ascending residue number, appending the one-letter amino-acid code only for residues that remained after trimming. Discontinuities in residue numbering or excised segments were encoded as ‘X’ placeholders, preserving registry with the original chain while masking non-overlapping regions. Provenance fields (superkingdom, UniProt accession, gene/protein symbol, organism and rank) were propagated into FASTA headers. For scalability, models were partitioned into balanced chunks and executed as an LSF job array; per chunk FASTA files were then concatenated to yield a single surface-trimmed sequence set for downstream alignment and ranking.Superkingdom partitioning and priority-preserving sequence clusteringThe surface-trimmed sequences set was partitioned by superkingdom from a single FASTA (plain or gzipped) using case-insensitive header classification. For each record, the first token before the first pipe (|) determined assignment to Archaea, Bacteria or Eukaryota; if no match was found, the full header was searched for canonical terms, otherwise the record was labelled Other. Records were streamed to four corresponding FASTA files and a brief count report, using a streaming approach to avoid loading the dataset into memory. We then generated a non-redundant representative set through an incremental, priority-aware MMseqs2 workflow: Archaea sequences were clustered first with linclust (minimum sequence identity 0.50; bidirectional coverage at least 0.50), establishing initial representatives; Bacteria were appended and incorporated with clusterupdate, which projects existing clusters onto the expanded database while preserving current representatives; the same append-and-update step was repeated for Eukaryota. Representatives from the final updated clustering were extracted and exported to FASTA, yielding a compact set that preserves the intended Archaea→Bacteria→Eukaryota priority.All-versus-all BLASTp on an LSF clusterUsing the representative FASTA produced above, we first created a BLASTp library (protein database) and then performed an all-versus-all BLASTp search by partitioning the FASTA into N whole-record chunks and submitting one job per chunk to the LSF scheduler (bsub). For each chunk, the BLASTp command targeted the newly built database and honoured user parameters (BLOSUM62 or BLOSUM45 with 14/2 gaps, E-value threshold, SEG filtering, xdrop and final xdrop, maximum targets and tabular outfmt 7 fields: qseqid pident length E-value bitscore stitle, with optional per job threads). Outputs were written as one TSV per chunk with corresponding logs, enabling straightforward monitoring and downstream collation of the all-versus-all results.BLASTp-derived protein similarity network constructionWe converted the per chunk BLASTp outputs (tabular outfmt 7) into a compact protein similarity network. BLASTp result lines were streamed, parsing query and subject identifiers from headers formatted as superkingdom|UniProt|gene|organism|rank. One node was created per unique UniProt accession with its associated annotations, and undirected edges were added between protein pairs supported by BLASTp matches. Self-hits were removed, and pair ordering was canonicalized to avoid duplicates. For each protein pair, several hits were aggregated to a single edge by retaining the most informative statistics (for example, maximum percent identity, minimum E-value, maximum bitscore and aligned length. The pipeline writes compressed edge and node tables (including each node’s unique-edge degree) and, optionally, a hits table, yielding an analysis-ready similarity network for downstream visualization and graph analysis.To define meaningful network connections, we applied calibrated similarity thresholds empirically in the BLASTp component. We used bitscore rather than E-value because bitscore is independent of database size, and retained edges with bitscore greater than or equal to 30. This threshold was chosen based on benchmarking against known GPCR relationships in the human GPCRome, which showed that higher cutoffs would exclude many established homologies. We also required at least 30% sequence identity to further minimize spurious links. As shown in Fig. 1g, although BLASTp detects local alignments, the aligned regions were distributed across the 7TM fold rather than confined to short local fragments, supporting broad structural coverage despite local sequence variability. To support reproducibility, we include the benchmarking analysis processed through our pipeline, along with the resulting Cytoscape network and the supplied network file.Seed-centric traversal of the similarity networkWe extracted seed-centred subnetworks from the BLASTp-derived graph using a deterministic go-forward procedure. Starting from the archaeal node with the highest unique-edge degree, candidates were evaluated in a fixed-priority order—higher percent identity, longer aligned length, higher bitscore, lower E-value—with ties resolved by rank (r1 > r2 > r3) and then by a stable identifier. Nodes passing preset quality gates were admitted once (visited-set enforcement), and expansion proceeded layer by layer until no additional qualifying neighbours remained or a predefined node/edge budget was reached. The method outputs Cytoscape-ready node and edge tables, along with summary metrics (subgraph size, degree distribution and threshold provenance).Strict reduction of the similarity network for CytoscapeWe transformed the BLASTp-derived protein similarity graph into a compact, high-confidence subnetwork suitable for Cytoscape by applying strict, staged reductions. Edge lists were cleaned to remove self-loops and duplicates, canonicalized to a single order per pair, and collapsed to one edge per protein pair by retaining the most informative statistics (for example, highest identity/bitscore, lowest E-value, longest aligned span). Edges were then filtered using stringent similarity gates (identity/bitscore/E-value and coverage), after which connected components were computed and constrained by global and per component node budgets. To control local density and highlight salient relationships, we limited each node to its strongest connections (top-k), enforced a minimum degree and iteratively pruned dangling nodes. The result is a Cytoscape-ready set of tables—reduced edges with summary metrics and reduced nodes with annotations and recomputed degrees—optimized for clear visualization and downstream analysis.Pairwise alignment and conservation mapping to PDB residue numberingPairwise sequence alignments were generated with Clustal Omega, and the alignment ‘traceback’ (symbol) line was parsed to classify each column as identity (*), conservative (:), semi-conservative (.) or gap. For every pair, we reconstructed a column-wise index map from alignment columns to the corresponding ungapped positions in each partner sequence, then translated those sequence positions to PDB residue identifiers (chain, residue number, insertion code) by loading the matching structures (PDB/mmCIF) and building sequence-to-structure mappings. The result was written as a per pair matches table (pos1(PDB) aa1 pos2(PDB) aa2 symbol) that preserves the original PDB residue numbering for both proteins, enabling structure-aware visualization and downstream analyses of conserved sites.Foldseek structural searches against the afdb50-minimal databaseWe queried individual protein structures against the prebuilt afdb50-minimal Foldseek database using foldseek easy-search. Before searches, the database prefix was normalized and, if needed, indexed. Runs shared the following settings: alignment type = local 3Di (–alignment-type 0), sensitivity = 9.5 (-s 9.5), coverage = 0.0 with cov-mode = query&target (-c 0.0–cov-mode 0), max-seqs = 10,000 (–max-seqs 10000), iterations = 1 (–num-iterations 1), exhaustive search on (–exhaustive-search), load-on-demand (–db-load-mode 2), threads = 12 (–threads 12) and graphics processing unit off. We then executed two representative calculations that differed only in E-value threshold: a permissive search (E = 1.0) and a stringent search (E = 0.01), producing separate.m8 outputs for downstream parsing.Automated pocket detection with pHinderWe ran pHinder in virtual screening mode on the AF2 model of TM184C (UniProt Q9NVA4, chain A), with only the virtual-screen branch enabled (virtualScreenSurfacesCalculation = 1; all other calculation toggles = 0). A high-resolution protein surface was computed and saved (circumsphere radius limit 6.5 Å, minimum patch area 10 Ų, high_resolution_surface = 1, save_surface = 1). A sampling grid was laid over the surface with grid_increment = 3.0 Å, sampling points were clash-filtered at 2.5 Å (virtual_clash_cutoff), and the remaining points were connected into a proximity graph with edges less than or equal to 2.0 Å (max_void_network_edge_length). Connected components with fewer than ten points (min_void_network_size) were discarded. For each surviving component, a triangulated void surface was generated and refined with one inward and one outward pass (IN: 1× at 2.0 Å; OUT: 1× at 2.0 Å). Network settings were left at defaults for this step (max_network_edge_length = 10.0 Å, min_network_size = 1, reduced representation and triangulation saving enabled). Hydrogens, waters and ions were excluded; logging was enabled; the Python recursion limit was set to 10,000; and execution used all central processing unit cores minus one. Outputs were written under the specified save path, yielding triangulated, pocket-shaped void surfaces suitable for downstream screening.Automated pocket detection with FpocketPutative ligand-binding pockets were identified with Fpocket, invoked by a lightweight Python wrapper in single-file mode on each PDB structure. The wrapper resolved the Fpocket executable from the system PATH, executed in the parent directory of PDB, captured return codes for provenance and collected results. For our runs, we specified a minimum pocket radius of 3.8 Å (flag -m 3.8); all other Fpocket options were left at their defaults (no condensed/energy modes, no ligand/chain restrictions and no overrides for clustering, distance metrics, α-sphere thresholds or grid settings). This provided reproducible pocket calls suitable for downstream aggregation and analysis.Upset plots for head-to-head comparison with FoldseekFor our head-to-head comparison with Foldseek, we used the test structures of rhodopsin (PDB 1F88, chain A), β2 adrenergic receptor (PDB 6KR8, chain A), adenosine receptor A2a (PDB 3VG9, chain A) and frizzled (PDB 8QW4, chain A), along with the search parameters described above. Structure-only results were summarized by TM-score and coverage (Extended Data Fig. 4a). For Extended Data Fig. 4b, we took the unique UniProt accessions returned for each query under three conditions—structure-only, Foldseek (E = 1.0), and Foldseek (E = 0.01)—with hits left geometrically unfiltered, constructed a simple membership matrix (IDs × queries) and computed set intersections (Supplementary Data 2). Upset grids display the largest intersections across the four queries (bar heights) alongside per-query totals (left bars), providing a direct, condition-by-condition view of shared versus query-specific hits.Superdark protein expressionUsing tissue- and cell-specific protein expression data downloaded as a flat file (.tsv) from the Human Protein Atlas (v.24, accessed 10 January 2025), we quantified superdark candidate expression patterns using custom Python code. We generated the heat maps in Extended Data Fig. 1a by assigning values of 1, 0.6, 0.3 and 0 to high, medium, low and not detected protein expression levels.Identifying arrestin codesUsing the built-in Python regular expressions library (re), we compiled regular expression patterns for short (r”[ST].[ST][^P][^P][STED]”) and long (r”[ST]..[ST][^P][^P][STED]”) arrestin codes. We then searched each superdark sequence for these string patterns, limiting our analysis to the C terminus as predicted in each AF2 model. All protein sequences were downloaded using the UniProt application programming interface in FASTA format.Identifying lysosome di-leucine motif codesWe used the same procedure for finding arrestin codes, but we used a compiled regular expression pattern for di_leucine motifs (r”[DE]…L[LI]”).Software usedThe following software was used: Conda (v.24.9.2), Python (v.3.8.17), Biopython (v.1.78), TM-align (v.20210224), Foldseek (v.10.941cd33), BLASTp (v.2.26.0+), Clustal Omega (v.1.2.4), pHinder (v.7.0), Fpocket (v.4.1), MMseqs2 (v.18.8cc5c), Cytoscape (v.3.10.3) and PyMOL (v.3.4.19).Cell cultureHEK293T and SK-MEL-28 cells were obtained from the American Type Culture Collection. HEK293A cells were obtained from Asuka Inoue—a base strain originally from Thermo Fisher. HEK293T and HEK293A cells were maintained in DMEM medium (Gibco) supplemented with 10% FBS and 1% penicillin–streptomycin. SK-MEL-28 cells were cultured in RPMI-1640 medium (Gibco) with 10% FBS and 1% penicillin–streptomycin. The hTERT human adult astrocytes were obtained from the Bayik laboratory (originally from University of California, San Francisco) and maintained in DMEM F12 with Glutagro (Corning), N-2 supplement (Thermo fisher), human epidermal growth factor (EGF) (PeproTech) 20 ng ml−1, human fibroblast growth factor (FGF) (PeproTech) 20 ng ml−1, 5% FBS, 1% penicillin–streptomycin. L1 cell line with mito-mCherry (from the Bayik laboratory) was maintained in DMEM F12 with Glutagro, B-27 supplement (Thermo fisher), EGF 20 ng ml−1, FGF 20 ng ml−1 and 1% penicillin–streptomycin. The primary GBM cell line (from the Ivan laboratory) was maintained in DMEM F12 with Glutagro, N-2 supplement, B-27 supplement, EGF 20 ng ml−1, FGF 20 ng ml−1, 1 mM sodium pyruvate, 2 μg ml−1 heparin (Sigma) and 1% penicillin–streptomycin in a flask with 5 μg cm−2 fibronectin (Corning). All cells were cultured in a humidified incubator at 37 °C with 5% CO2. The passage number was recorded for each cell line.Lentivirus productionLentiviral particles were generated using HEK293T cells. Cells were seeded at a density determined by the size of the dish or plate in DMEM containing 10% FBS and 1% penicillin/streptomycin. After overnight incubation, cells were transfected using PolyJet (SignaGen Laboratories) at a 1:3.33 plasmid:PolyJet ratio, with a plasmid mixture containing a 4:3:1 ratio of transfer:Δ8.2:vsv-g plasmids. Approximately 48 h after transfection, culture media were collected and replaced with fresh media. Viral supernatant was collected on days 2 and 3 following transfection. The collected supernatants were combined and filtered through a 0.45-μm mixed cellulose esters filter (Millipore) to remove cellular debris, and the filtrate was stored at −80 °C until further use.Stable cell line generation using lentiviral transduction for imagingAll stable cell lines with TM184A, TM184B and TM184C full-length and mutants tagged with eGFP or mCherry were generated for each cell line in the same way, except each had the appropriate concentration of puromycin (HEK293A, SK-MEL-28, HFF-1, primary GBM cell line = 1 μg ml−1). Cells were seeded at 50–70% confluency per well in a six-well plate. The following day, cells were infected with filtered lentiviral particles. After 24 h, the virus-containing medium was removed and replaced with fresh medium. The next day, cells were split, maintaining a confluence of 80–100%, and, appropriate to the cell line, the medium was added. There was always a non-transduced control for selection to ensure the puromycin concentration was kept at the lowest possible level. After the cells recovered, they were further split into larger dishes and maintained at 80–100% confluence. The hTERT human astrocytes and L1 mito-mCherry PDX cell line were selected using fluorescence-activated cell sorting (FACS) on a BD FACSAria II (BD Biosciences) with FACSDiva software instead of puromycin. In brief, either the TM184C–eGFP-positive or the double-positive population of L1 mito-mCherry and TM184C–eGFP was sorted into 12-well or 24-well dishes (Supplementary Fig. 2). All stable cell lines were maintained as polyclonal populations at the appropriate puromycin concentration. HEK293T and SK-MEL-28 cells (American Type Culture Collection) are authenticated by the vendor at source; the HEK293A, HFF-1, hTERT-immortalized astrocyte and patient-derived (L1 xenograft and primary GBM) lines were not re-authenticated independently by short tandem repeat profiling. All cell lines tested negative for mycoplasma contamination by PCR. None of the lines used (HEK293T, HEK293A, SK-MEL-28, HFF-1, hTERT astrocytes, L1, primary GBM) appear on the ICLAC Register of Commonly Misidentified Cell Lines.For co-localization experiments, the stable cell lines with full-length or mutant TM184C tagged with eGFP or mCherry were seeded in confocal dishes to reach 30–50% confluency the next day. For imaging astrocytes and primary GBM cell lines, the plates were coated additionally with 5 μg ml−1 fibronectin for 30–60 min. Then these were transfected with TransiT-2020 (3 μl:1 μg of plasmid) and 1 μg of the indicated plasmid, each tagged with mVenus, mCherry, mRFP1 or YFP. The cells were imaged 24–48 h after transfection, and treatments with autophagy modulators (‘Autophagic flux’) were performed for the indicated time period and dosage, coinciding with the day of imaging.PlasmidsARRB1- and ARRB2-Venus were gifts from K. D. G. Pfleger. We took arestin1 and 2 to create 3xHA–ARRB1 and 3xHA–ARRB2. The plasmids containing all mini G, 4AG alpha and bystander markers with mVenus or RLuc8 were gifts from N. Lambert. pLJM1–eGFP, pLJM1-LAMP1-mRFP1-FLAG, pcDNA3.1-mCherry-hLC3B, lentiCRISPR-v2, lentiCRISPR-v2 blast, mCherry-ActA-IRES-puromycin-pLVx-EF1a, TRUPATH plasmid kit and PRESTO-Tango plasmid kit were all obtained from Addgene.pcDNA3.1(+)-Superdark constructs were generated using HiFi DNA assembly (New England Biolabs). Superdark genes were codon-optimized for humans (GeneScript), then obtained as gBlocks (Integrated DNA Technologies) and amplified by PCR using primers designed to overlap with the pcDNA3.1(+) backbone and the start and stop codons of each superdark sequence. Backbone and insert constructs were co-incubated with HiFi master mix and transformed into competent DH5α Escherichia coli.The pcDNA3.1(+)-Superdark-RLuc8 constructs were generated using HiFi DNA assembly. The linker (GVPRARDPPVAT) and RLuc8 tag were amplified from GPR4-RLuc8 (gift from N. Lambert), using primers that provided homology to the C-terminal tails of the superdark constructs and pcDNA backbone. The backbone with superdark gene and RLuc8 insert constructs were co-incubated with HiFi master mix and transformed into competent DH5α cells to form fully intact plasmids.For adding or changing small DNA sequences (HiBiT and HA tags, point mutations, gRNA sequences), the same strategy, called 5’-phosphate PCR assembly, was used. Briefly, vectors were amplified using Q5 High-Fidelity DNA Polymerase (NEB) reactions with a 5’-phosphate-modified complementary primer, a primer containing the desired change, a complementary region (in opposite orientations) and the template plasmid DNA. PCR products were verified by agarose gel electrophoresis, then digested with DpnI (NEB) at 37 °C for 1.5 h, followed by heat inactivation. Ligation was performed overnight at 16 °C or for 2 h at room temperature using T4 DNA ligase (NEB). Ligase was heat-inactivated, and 2 µl of the ligation mix was transformed into competent DH5α cells.All plasmids were verified by Sanger or nanopore sequencing (Eurofins Genomics).CRISPR–Cas9-mediated genome editingPlasmidsPlasmids were constructed using the above 5′-phosphate PCR assembly strategy. Six gRNA sequences targeting the TM184C mRNA were selected and screened for optimal efficiency. Two of the six gRNA sequences are used and shown in the study.Transient transfection in HEKHEK293T cells were seeded at 300,000 cells per well in six-well plates. After 24 h, cells were transfected transiently with 1.25 μg of plasmid DNA and 3 μl μg−1 of POLYJET in Advanced DMEM (Gibco). After 24–48 h, the cells were trypsinized and resuspended in medium containing puromycin at 1 μg ml−1. Cells were split into the desired-sized plate or dish. There was always a non-transfected control for the selection to ensure the puromycin selection agent was working as expected. Once the controls were dead, the experimental cells were then collected for analysis.Rescue experimentsBecause TM184C behaved as an essential gene, the established stable cell lines with full-length or mutant TM184C tagged with eGFP or mCherry, and parental HEK293A cells, were used as the base cells. The stably integrated GFP-tagged transgenes are codon-optimized and therefore resistant to the endogenous TM184C-targeting gRNAs. In brief, they were transfected with lentiCRISPR-v2 blast plasmids with corresponding gRNAs targeting TM184C, then selected with puromycin at 1 μg ml−1 and blasticidin at 10 μg ml−1.HiBiT LgBiT assayHEK293T cells were seeded at 300,000 cells per well in six-well plates. After 24 h, cells were transfected transiently with 1 μg of N-terminal tagged HiBiT receptor, 3 μl of TransIT-2020 (Mirus) transfection reagent per microgram of DNA and 100 μl of advanced DMEM (ADMEM). Control samples were transfected with empty vector. Cells were washed with PBS as described in the bystander BRET protocol and resuspended before transferring 30 μl into a 384-well plate in quadruplicate. A 1:100 dilution of LgBiT and a 1:50 dilution of substrate solution (Promega, Nano-Glo HiBiT Extracellular Detection Kit) were prepared and mixed. A 30-μl aliquot of this mixture was transferred to a 384-well plate, and luminescence was measured after a 10-min incubation using a ClarioStar plate reader. Emission was read at 450-25, gain 3,600, with a measurement interval time of 0.10 s.Bystander BRET assayHEK293T cells were seeded at 300,000 cells per well in six-well plates. After 24 h, cells were transfected transiently with 200 ng of RLuc8-tagged receptor, 1 μg of mVenus-labelled bystander protein or empty vector for RLuc8 control, and 3 μl of TransIT-2020 (Mirus) per microgram of DNA in 120 μl of ADMEM. After 48 h, the medium was aspirated, and the cells were resuspended in 1 ml of PBS. A 300-μl aliquot was transferred to a 96-well plate and centrifuged at 300 relative centrifugal force (RCF) for 3 min. The supernatant was aspirated, and 150 μl of fresh PBS was added, followed by gentle resuspension of the pellet. A 30-μl aliquot of the cell suspension was transferred into a 384-well white-bottom plate in quadruplicate. Coelenterazine-h (Nanolight Technology) was added at a final concentration of 5 μM for a total volume of 45 μl. The plate was vortexed for 15 s at 1,750 rpm before luminescence readings were taken using a ClarioStar plate reader in multichromatic mode, with emissions 532.5-25 and 485-20, gain 3,600 and a measurement interval time of 1 s. Emission intensities for RLuc8 and Venus were recorded, and the net BRET ratio was calculated by subtracting the RLuc8–Venus ratio from the RLuc8 control.4AG protein recruitment assayFor 4AG protein coupling, HEK293T cells were seeded at 300,000 cells per well in six-well plates. After 24 h, cells were transfected with Receptor–RLuc8, 4A subunit, Venus-1–155-Gγ2, Venus-155–239-Gβ1 and pcDNA3.1(+) in a (1:10:5:5:5) ratio for a total of 2.6 μg of plasmid DNA and 3 µl µg−1 TransIT-2020 (Mirus) in each well of a six-well plate. After a 24–48-h incubation following transfection, the plates were aspirated, cells were washed, resuspended and 30 μl was added to a 384-well plate in quadruplicate, as in the bystander BRET protocol. Coelenterazine-h was then added at a 5 μM final concentration for a final volume of 45 µl. The plate was then read using the ClarioStar settings from Bystander BRET.TRUPATH G-protein dissociation assayHEK293T cells were seeded at 300,000 cells per well in six-well plates. After 24 h, cells were transfected transiently using a 1:1:1:1 DNA ratio of receptor: Gα–RLuc8:Gβ:Gγ–GFP2 (100 ng per construct for six-well dishes). TransIT-2020 (Mirus) was used to complex the DNA at a ratio of 3 µl TransIT DNA per microgram DNA in ADMEM. At 48 h after transfection, the plates were aspirated, cells were washed, resuspended and 30 μl was added to a 384-well plate in quadruplicate, as in the bystander BRET protocol. Prolume Purple (Nanolight Technology) was then added to a final concentration of 5 μM, resulting in a final volume of 45 µl. The plate was then read on the ClarioStar in multichromatic mode at emissions 400-10 and 520-10, with a 3,600 gain, spiral average, 2-mm diameter and a 0.27 s measurement interval. The same process was repeated, but with varying receptor concentrations to increase constitutive activity through Gs dissociation compared with DRD2.β-Arrestin-BRET recruitment assayHEK293T cells were seeded at 300,000 cells per well in six-well plates. After 24 h, cells were transfected transiently with a 1:10 ratio of receptor-RLuc8: Arrestin-mVenus (1.1 μg total DNA), using 3 μl of TransIT-2020 (Mirus) transfection reagent per microgram of DNA in 160 μl of ADMEM. After 48 h, the protocol was identical to the bystander BRET assay described above. The plate was then read using the ClarioStar settings from Bystander BRET.β-Arrestin-BRET and GRK recruitment assayHEK293T cells were seeded at 300,000 cells per well in six-well plates. After 24 h, cells were transfected transiently with a 1:5:10 ratio of receptor-RLuc8: GRK (2,3,5,6) or empty vector: Arrestin-mVenus (1.6 μg total DNA) using 3 μl of TransIT-2020 (Mirus) transfection reagent per microgram of DNA in 160 μl of ADMEM. After 48 h, the protocol was identical to the bystander BRET assay described above. The plate was then read using the ClarioStar settings from Bystander BRET.GRK2 recruitment assayHEK293T cells were seeded at 300,000 cells per well in six-well plates. After 24 h, cells were transfected transiently with a 1:10 ratio of receptor-RLuc8: GRK2-mVenus (1.1 μg total DNA), using 3 μl of TransIT-2020 (Mirus) transfection reagent per microgram of DNA in 160 μl of ADMEM. After 48 h, the protocol was identical to the bystander BRET assay described above. The plate was then read using the ClarioStar settings from Bystander BRET.ImmunoprecipitationHEK293 T cells were seeded at 3 million cells in a 10-cm plate. The following day, cells were transfected transiently with 4 μg of DNA for single endogenous IPs, 2 μg:2 μg for two-construct Co-IPs or 3 μg:1 μg for Co-IP with receptors and β-arrestins, using 3 μl of TransIT transfection reagent per microgram of DNA in 400 μl of ADMEM. After 48 h, the medium was aspirated, and cells were washed with cold PBS and centrifuged at 1,000 RCF for 5 min. The PBS was aspirated, and cells were lysed using lysis buffer containing 20 mM HEPES, pH 7.5, 100 mM NaCl, 1 mM EDTA, 1 mM phenylmethylsulfonyl fluoride (PMSF), 1× EDTA-free protease inhibitor cocktail (Sigma), 1 mM MgCl2, 10 mM β-glycerophosphate (Sigma), 10 mM sodium pyrophosphate (Sigma) and 1% Triton X-100 (Sigma). Lysis was performed for 30 min with vortexing every 5 min. The lysate was centrifuged at 14,000 RCF for 10 min; 500 μl of the supernatant was combined with 10 μg of anti-FLAG M2 antibody (Sigma, catalogue no. F1804) or anti-HA.11 antibody (Biolegend, catalogue no. 901502) and incubated for 1–2 h or overnight at 4 °C with end-over-end mixing.Pierce Protein A/G Magnetic Beads (0.25 mg; Thermo Fisher) were added to a 1.5-ml microcentrifuge tube for immunoprecipitation. Beads were washed with 175 μl of Wash Buffer (TBS, 0.05% Tween-20 and 0.15 M NaCl) and vortexed gently before collecting using a magnetic stand; the supernatant was discarded. The beads were washed three times with the wash buffer. The antigen–antibody mixture was then added to the beads and incubated at room temperature for 1 h with gentle mixing. The beads were collected using a magnetic stand and washed twice with 500 μl of wash buffer, then washed a final time with 500 μl of purified water. The beads were collected on a magnetic stand and the supernatant was discarded.For elution, 50 μl of low-pH elution buffer (0.1 M glycine, pH 2.0) was added, and the beads were incubated at room temperature with mixing at 1,400 rpm for 10 min. Beads were collected using a magnetic stand, and the supernatant was transferred to a new tube containing 7.5 μl of neutralization buffer (1 M Tris, pH 9.0) per 50 μl of eluate. The final samples were analysed using SDS–PAGE and western blot.Immunoblotting from cell linesCells were disrupted on ice using the same lysis buffer as the one above. The resulting lysates underwent centrifugation to remove cellular debris, and protein concentrations were quantified using the BCA Protein Assay (Thermo Fisher Scientific). Samples were then separated by electrophoresis and transferred onto 0.22-μm polyvinylidene fluoride membranes (GenScript). The membranes were blocked in Tris-buffered saline Tween (TBST) containing 5% milk (ApexBio) before incubation with the following primary antibodies: Histone H3 (Cell Signaling Technology (CST), catalogue no. 12648), FLAG M2 (Sigma, catalogue no. F1804), TM184C (Atlas Antibodies, catalogue no. HPA054013), LC3B (CST Autophagy Atg8 Family Antibody Sampler Kit, catalogue no. 64459) and VDAC1/Porin (Abcam, catalogue no. ab110326, clone 16G9E6BC4; yeast loading control). Following washes in TBST, membranes were incubated with the following secondary antibodies: anti-rabbit IgG HRP conjugate (Cytiva, catalogue no. NA934) and anti-mouse IgG HRP conjugate (Cytiva, catalogue no. NA931). After further washes with TBST, protein bands were visualized through chemiluminescence using Clarity Western ECL substrate (Bio-Rad) or diluted SuperSignal West Atto substrates (Thermo Fisher Scientific). The densitometric analysis of the blots was performed using ImageJ. Fold changes in expression levels of the construct were analysed relative to the corresponding control signals.Immunoblotting from yeastYeast cells were grown overnight in YPD medium. Cells equivalent to an optical density (OD)600 of 1.2 were collected by centrifugation (6,000 RCF, 1 min) and washed with sterile water. The cell pellet was resuspended in 75 µl of Rodel Mix (0.37 M NaOH, 8.9% v/v β-mercaptoethanol, and 10 mM PMSF added fresh), then 500 µl of sterile water was added. After vortexing, an equal volume of 50% trichloroacetic acid was added. The mixture was incubated on ice for 10–15 min and centrifuged (14,000 RCF, 4 °C, 10 min). The supernatant was discarded, and the pellet washed sequentially with 500 µl of 0.5 M Tris base (unadjusted pH) and 500 µl of sterile water (14,000 RCF, 4 °C, 5 min) without disturbing the pellet. The final pellet was resuspended in 25 µl of resuspension buffer (100 mM NaCl, 20 mM HEPES, pH 7.4, 1 mM EDTA, 1 mM fresh PMSF), and the protein concentration was measured by BCA assay. The SDS sample buffer was loaded onto the SDS–PAGE gel for further analysis.Autophagic fluxTo measure autophagic flux upon knockdown of TM184C, the CRISPR-edited HEK293T cells selected with puromycin (see ‘CRISPR–Cas9-mediated genome editing’) were treated with 1 nM of Concanamycin A (Cayman Chemicals) or vehicle (0.06% acetonitrile final) for 24 h. The cells were then processed for LC3B immunoblotting. For imaging, Concanamycin A or Bafilomycin A1 (CST) at 1 nM was used for 18–24 h, as indicated in the legends.Protein expression and purificationHEK293T cells expressing FLAG-tagged TM184C were collected by scraping, washed with cold PBS, then centrifuged and flash-frozen in liquid nitrogen. For membrane preparation, thawed cells were homogenized in low-salt buffer (10 mM HEPES, pH 7.5, 10 mM MgCl2, 20 mM KCl) supplemented with protease inhibitors (Roche). The homogenate was incubated for 15 min at room temperature with benzonase (Millipore) and MgCl2 to a final concentration of 2.5 mM. Following homogenization, cell lysates were subjected to ultracentrifugation at 105,000 RCF for 40 min at 4 °C to isolate membranes. This was repeated twice with a high-salt buffer (10 mM HEPES, pH 7.5, 10 mM MgCl2, 20 mM KCl, 1 M NaCl) supplemented with protease inhibitors, and the membranes were then ready for downstream purification. They were solubilized in 1%/0.2% dodecyl-β-d-maltoside (DDM)/ cholesteryl hemisuccinate (CHS) (Anatrace) HNG buffer (20 mM HEPES pH 7.5, 300 mM NaCl, 10% glycerol) with 1 mM EDTA for 3 h with rotation at 4°C, clarified by ultracentrifugation at 105,000 RCF for 40 min at 4 °C, then incubated with washed anti-FLAG M2 resin (Sigma) overnight. The beads were then collected in a column and washed three times with ten column volumes of 0.1%/0.02%, 0.05%/0.025% and 0.025%/0.0125% DDM/CHS HNG buffer. FLAG–TM184C was eluted in flag elution buffer (300 mM NaCl, HNG buffer 1 mM EDTA, 0.025%/0.0125% DDM/CHS + 200 μg ml−1 FLAG peptide). The eluates were then concentrated in a Vivaspin 500 (Cytiva) with a 100,000 molecular weight cut-off and loaded onto an SDS–PAGE gel for further analysis. All buffers were prepared with 0.22-µm-filtered reagents, and all steps were performed at 4°C or on ice.Lambda phosphatase treatment of cell lysatesHEK293T lysates expressing FLAG–TM184C and empty vector, GRK2, GRK3, GRK5 and GRK6 were prepared using a lysis buffer consisting of 20 mM HEPES pH 7.5, 100 mM NaCl, 1 mM EDTA, 1% Triton X-100, 1× protease inhibitor cocktail and 1 mM PMSF. Following a 20-min incubation on ice, lysates were centrifuged at 14,000 RCF for 10 min. A 50 µl aliquot of the supernatant was transferred to a PCR tube and 5 µl of 10× MnCl2 (NEB) was added, followed by 2.5 µl of lambda phosphatase (NEB). The reaction was incubated at 30 °C for 1 h. The addition of SDS sample buffer terminated the reaction, and samples were incubated at room temperature for 5 min before loading onto a gel for further analysis.Yeast strainsYeast strains were maintained and grown in a YPD medium composed of 1 g l−1 yeast extract, 2 g l−1 peptone and 20 g l−1 dextrose. Synthetic complete medium (SCM) is composed of 1.7 g l−1 yeast nitrogen base without amino acids and ammonium sulfate, 20 g l−1 dextrose, 5 g l−1 ammonium sulfate, 20 g l−1 d-glucose and the recommended amount of complete supplement mixture—complete or without uracil powder (MP Biomedicals).gRNA plasmid and hfl1Δ strain generationA CRISPR gRNA plasmid targeting HFL1 (pML104-HFL1.1031) was constructed in the pML104 vector using the 5’ phosphate PCR assembly. A 100-bp repair DNA payload for hfl1Δ was created by annealing complementary 50-bp homology arm oligonucleotides designed to flank the HFL1 coding region. Annealing was verified by agarose gel electrophoresis. S. cerevisiae BY4741 strains containing a pre-integrated X-2 landing pad (BY4741 X-2 LP) were transformed with pML104-HFL1.1031 and the repair DNA payload using a standard lithium acetate/polyethylene glycol method. The hfl1Δ yeast strain was verified by the following method. Yeast gDNA was extracted from transformants and PCR was performed using primers flanking the HFL1 coding region. hfl1Δ deletion was confirmed by agarose gel electrophoresis based on PCR product size. Verified hfl1Δ strains were counter-selected on 5-fluoroorotic acid plates to remove the CRISPR plasmid.Superdark protein integration into hfl1Δ X-2 LP strainsFor integration of superdark proteins (TM184A, TM184B, TM184C) or HFL1 into the X-2 LP of hfl1Δ strains, a PCR-generated DNA payload containing superdark protein sequences or the HFL1 gene and X-2 LP homology arms using either the pcDNA3.1 plasmids or the yeast genome as a template. hfl1Δ strains were transformed with pML104-X2 gRNA plasmid (targeting the X-2 LP) and the DNA payload using the same lithium acetate/polyethylene glycol method as above.FM4-64 staining of yeast vacuolesStrains were grown in SCM overnight at 30 °C with shaking. The following day, yeast cells were diluted to an OD = 0.2 in 5 ml of SCM and grown in a shaking incubator at 30 °C for around 36–40 h, reaching saturation. Finally, yeast cells were set to an OD of 1.0 in fresh SCM for confocal analysis. A 1-ml aliquot of yeast cells, normalized to an OD of 1.0, was transferred to a 1.5-ml sterile Eppendorf tube and centrifuged at 10,000 RCF for 3 min. The supernatant was removed, and yeast cells were resuspended in 100 μl of fresh SCM. Then, 3 μl of 1 mM FM4-64 dye was added to the sample, vortexed and incubated at 30 °C (no shaking) for 20 min (final FM4-64 dye concentration around 29.1 μM). Cells were centrifuged at 10,000 RCF for 3 min, and the supernatant was discarded. Yeast cells were washed with 100 μl of fresh SCM, centrifuged at 10,000 RCF for 3 min and the supernatant was discarded. This process was repeated for a total of two washes. Yeast cells were resuspended in 200 μl of fresh SCM before 6 μl of the yeast mixture was plated on a 2% SCM agarose pad, and a glass coverslip was dropped onto the agarose pad.MS data acquisition and analysis from on-bead digestionTwo immunoprecipitation samples were subjected to on-bead tryptic digestion according to an established protocol53. Bead-bound proteins were digested by adding 10 µl of trypsin (10 ng µl−1) in 100 mM ammonium bicarbonate (AMBIC) at an enzyme-to-protein ratio of 1:100 (wt/wt). Beads were vortexed every 2–3 min for the first 15 min to ensure homogeneous protease distribution, then incubated overnight at 37 °C. A second aliquot of trypsin was added the following day for an additional 4-h digestion at 37 °C. Supernatants were collected using a magnetic rack, acidified to 5% formic acid (v/v), and desalted using C18 ultra micro spin columns per the manufacturer’s instructions. Samples were dried by vacuum centrifugation and reconstituted in 1% acetic acid before analysis.Peptides were separated and analysed on a ThermoScientific Fusion Lumos Tribrid mass spectrometer coupled to a Dionex Ultimate 3000 RSLCnano HPLC system. Each sample (5 µl) was loaded onto an Acclaim PepMap C18 trapping column (100 µm × 2 cm, 5 µm, 100 Å), then resolved on an Acclaim PepMap C18 analytical column (75 µm × 25 cm, 2 µm, 100 Å) using a 140-min gradient from 2% to 90% acetonitrile in 0.1% formic acid at 0.3 µl min−1. The microelectrospray ion source was operated at 2.5 kV. Data-dependent acquisition was performed with a 3-s duty cycle: full MS1 scans were acquired at a resolution of 120,000 over m/z 350–1,500 (automatic gain control target 4.0 × 105), followed by MS2 fragmentation using collision-induced dissociation (35% collision energy, 1.6 Da isolation window, ion trap detection, automatic gain control 2.0 × 10³). Dynamic exclusion was applied for 60 s within a 10 ppm window after one repeat.Raw data were searched against the human SwissProtKB database (26,576 entries, downloaded March 2022) using Sequest within Proteome Discoverer v.2.5. Search parameters included full tryptic cleavage with up to two missed cleavages, 10 ppm MS1 and 0.06 Da MS2 mass tolerances, variable oxidation of methionine and N-terminal acetylation, and static carbamidomethylation of cysteine. Peptide and protein identifications were validated using Percolator with a reverse-decoy strategy, requiring at least two peptides per protein and a false discovery rate less than or equal to 1% at both the peptide and protein levels. To investigate post-translational modifications, the dataset was also queried against the TM184C sequence with phosphorylation included as a variable modification, yielding three phosphopeptides with distinct phosphorylation sites mapping to serine residues S422, S432 and S435.MS data acquisition and analysis from gelPurified TM184C was loaded onto an SDS–polyacrylamide gel and stained with a Coomassie dye. The band indicated in the figure was cut to minimize excess polyacrylamide and divided into several smaller pieces for protein digestion. The gel pieces were washed with water and dehydrated in acetonitrile. The bands were then reduced with dithiothreitol and alkylated with iodoacetamide before the in-gel digestion. All bands were digested in-gel using trypsin by adding 5 μl 10 ng μl−1 trypsin or chymotrypsin in 50 mM ammonium bicarbonate and incubating overnight at room temperature to achieve complete digestion. The peptides formed were extracted from the polyacrylamide in two aliquots of 30 µl each, using 50% acetonitrile with 5% formic acid. These extracts were combined and evaporated to less than 10 μl in a Speedvac, then resuspended in 1% acetic acid to a final volume of approximately 30 μl for liquid chromatography–mass spectrometry analysis.The liquid chromatography–mass spectrometry system was a Bruker TimsTof Pro2 Q-Tof mass spectrometer operating in positive-ion mode, coupled with a CaptiveSpray ion source (Bruker Daltonik GmbH). The HPLC column was a Bruker 15 cm × 75 µm id C18 ReproSil AQ, 1.9 μm, 120 Å reversed-phase capillary chromatography column. Extract (1 μl) was injected, and the peptides eluted from the column by an acetonitrile/0.1% formic acid gradient at a flow rate of 0.3 μl min−1 were introduced into the mass spectrometer source online. The digests were analysed using a parallel accumulation serial fragmentation DDA method to select precursor ions for fragmentation with a trapped ion mobility–mass spectrometry scan followed by ten parallel accumulation serial fragmentation tandem mass spectrometry scans. The trapped ion mobility–mass spectrometry survey scan was acquired over 0.60–1.6 V cm−2 and 100–1,700 m/z, with a ramp time of 166 ms. The total cycle time for the parallel accumulation serial fragmentation scans was 1.2 s, and the tandem mass spectrometry experiments were performed with collision energies between 20 eV (0.6 V cm−2) and 59 eV (1.6 V cm−2). Precursors with two to five charges were selected with the target value set to 20,000 a.u. and an intensity threshold of 2,500 a.u. Precursors were dynamically excluded for 0.4 s.Data were analysed using all collision-induced dissociation spectra collected in the experiment to search the human SwissProt database using the program MSFragger. These samples were analysed using a low-coverage gradient from 2% to 70% acetonitrile over 110 min.Confocal microscopyMicroscopy was performed using a Leica Stellaris 5 confocal inverted microscope equipped with a ×63 oil-immersion objective lens and operated by Leica Application Suite X software (LAS X; v1.4.6.28433). Image acquisition was performed by a line-scan. The Leica LAS X Dye Assistant module was used to obtain optimal spectral yield and minimal overlap. The acquisition format was 1,024 × 1,024, speed 400–700, bidirectional X and pinhole at Airy 1. The four HyD detectors were set to analogue or counting mode with 16-bit depth for vesicle transfer and immunofluorescence experiments. All mammalian cells were seeded onto 35-mm confocal dishes with 20-mm glass bottoms (VWR, catalogue no. 75856-742) or in Nunc 27 mm glass-bottom dishes (Thermo Fisher Scientific, catalogue no. 150682) and cultured for 24 h or 48 h under standard conditions (37 °C, 5% CO2). For immunofluorescence, the same dishes and culture conditions were used, followed by standard fixation and the antibody protocol described in ‘Fixed-cell immunofluorescence microscopy sample preparation’. For fixed mammalian cell samples, images consisted of z-stacks determined by the Leica LAS X System Optimized Z sectional thickness. For live-cell mammalian imaging experiments, three to five optical sections per z-stack were acquired, using the time interval minimize feature.Fixed-cell immunofluorescence microscopy sample preparationFor immunofluorescence microscopy, cells were fixed in 4% paraformaldehyde for 10 min at room temperature, followed by three 5-min washes in PBS. Cells were permeabilized using 0.1% Triton X-100 in PBS for 10 min, followed by an additional PBS wash. Blocking was performed with 1% bovine serum albumin in PBS for 1 h at room temperature. Primary antibody incubation was carried out using rabbit anti-human TM184C antibody (Thomas Scientific, catalogue no. HPA054013-100) on wild-type HEK293A cells. For TM184C–eGFP fusion stable line immunofluorescence, KIF5B (CST, catalogue no. 62696) primary antibody was used. All antibodies were diluted in blocking buffer overnight (12 h) at 4 °C. After three PBS washes, cells were incubated with goat anti-rabbit IgG conjugated to Alexa Fluor 647 (CST, catalogue no. 4414S) for 1 h at room temperature. Following final PBS washes, samples were mounted using ProLong Diamond Antifade Mountant containing DAPI (Thermo Fisher Scientific, catalogue no. P36962).Live-cell microscopyAll mammalian cells were seeded onto 35-mm confocal dishes with 20 mm glass bottoms (VWR, catalogue no. 75856-742) or in Nunc 27 mm glass-bottom dishes (Thermo Fisher Scientific, catalogue no. 150682) and cultured for 24 h or 48 h. For live-cell imaging, the culture medium was replaced with FluoroBrite DMEM (Thermo Fisher Scientific, catalogue no. A1896701). Depending on the experimental conditions, cells were treated with one or more of the following fluorescent markers: NucBlue Live ReadyProbes Reagent (Thermo Fisher Scientific, catalogue no. R37605), Tubulin Tracker Deep Red (Thermo Fisher Scientific, catalogue no. T34076) or LysoTracker Red DND-99 (Thermo Fisher Scientific, catalogue no. L7528). Each dye was incubated for 15 min, followed by exchange into fresh FluoroBrite DMEM.Imaging of yeast vacuolesYeast vacuoles were stained using FM4-64 dye (Thermo Fisher Scientific). FM4-64 powder was dissolved in nuclease-free water to a stock concentration of 1 mM and stored in aliquots protected from light at −20 °C. For staining, yeast cells normalized to OD600 = 1.0 were pelleted by centrifugation at 10,000g for 3 min and resuspended in 100 μl fresh SCM containing FM4-64 dye at a final concentration of 29 μM. After incubation for 20 min at 30 °C without shaking, cells were washed twice by centrifugation at 10,000g for 3 min each and resuspended in fresh SCM.FM4-64-stained yeast cells (6 μl) were pipetted gently onto a microscope slide containing a freshly prepared pad of solidified agarose (2%) dissolved in SCM. Cells were distributed evenly across the agarose pad surface and allowed to partially dry before a glass coverslip was placed on top. Coverslips were secured to slides using nail polish applied to all four corners, which was allowed to dry completely before imaging.Yeast cell images were acquired similarly as z-stacks optimized for visualization of vacuoles labelled with FM4-64. Image processing was performed using the Leica Lightning module, followed by maximum-intensity projection or 3D visualization, as appropriate.Automated quantification of co-localization (vesicle co-localization software)To quantify spatial overlap between vesicle populations, two differentially coloured fluorescence channels (Channel 1 and Channel 2) were processed using a custom Python application built with OpenCV and NumPy. Fluorescence images were either opened directly in greyscale or converted to greyscale during preprocessing. The resulting single-channel images were then used for all downstream analysis. Each channel was thresholded at three intensity levels (minimum, intermediate and maximum) to define binary masks representing signal-positive pixels. Contours were extracted from the thresholded regions, filtered by area to remove noise and artifacts and combined into a single mask for each channel.After mask generation, pixel-level co-localization was calculated as the logical intersection between the two masks using a bitwise ‘AND’ operation. For each channel, the total number of signal pixels and the number of overlapping pixels were counted, yielding both absolute overlap counts and fractional overlaps relative to each channel (fraction = overlap/total pixels). These metrics provided a quantitative estimate of co-localized fluorescence.Colour overlays highlighting overlapping regions were generated for visualization, and all images (masked, contour-only, overlay and overlap-highlighted) were saved to disk. A time-stamped log file automatically recorded image file paths, threshold and contour parameters, total pixel counts, overlap counts and fractional co-localization values for full reproducibility. All accompanying code is available on GitHub.Automated bpp score softwareTo quantify intercellular connectivity, bpp scores were computed from fluorescence images using a custom Python application built with OpenCV and NumPy. Images were first converted to greyscale if needed and smoothed with a low-pass filter to reduce pixel-scale noise. Background was then suppressed using intensity thresholding and morphological opening, yielding a binary mask of TM184C-positive structures.From this mask, contours were extracted and converted into skeletal centre lines using morphological thinning. The skeletonized network was parsed into discrete segments, and endpoints, branch points and segment lengths were computed. Segments connecting two distinct cell bodies were classified as bridges, elongated segments extending from a single-cell body as projections, and shorter or less elongated segments as protrusions based on empirically defined length and shape criteria.Segments connecting two distinct cell bodies were classified as bridges, elongated segments extending from a single-cell body as projections and shorter or less elongated segments as protrusions, based on empirically defined length and shape criteria. For each field of view, the bpp score was then calculated as the sum of weighted contributions from bridges, projections and protrusions according to$${\rm{bpp}}=\sum _{i\in B}\left[({n}_{c,i}-1)\left(10+\frac{{d}_{{\rm{mm}},i}}{2}\right)\right]+\sum _{i\in J}\left[0.5\left(\frac{{d}_{{\rm{mm}},i}}{5}\right)\right]+\sum _{i\in T}[0.05],$$where \(B\) denotes bridge segments, \(J\) junctional projections, \(T\) protrusions, \({n}_{c,i}\) the number of connected cells for a bridge \(i\) and \({d}_{\text{mm},i}\) its length in millimetres. These scores were averaged across images and biological replicates to quantify TM184C-dependent changes in intercellular connectivity. All accompanying code is available on GitHub.Automated vesicle transfer detectionTo quantify intercellular vesicle exchange using vesicle triangulator software, we developed a custom Python application built with OpenCV, NumPy and a computational geometry library to analyse two-colour TM184C–eGFP and TM184C–mCherry images. Fluorescence images for each channel were opened or converted to greyscale, and TM184C-positive vesicles were segmented using multilevel intensity thresholding and connected-components analysis with area filtering to remove noise and non-vesicular objects. For each channel, vesicle centroids were extracted and saved as two-dimensional point clouds in CSV format, preserving each vesicle’s colour identity (donor versus recipient).Point clouds from both channels were pooled and used to construct a two-dimensional Delaunay triangulation, treating vesicle centroids as vertices in a planar graph. Edges longer than a user-defined cut-off were removed to restrict analysis to local vesicle neighbourhoods, and remaining edges were classified as same-colour or mixed-colour based on the labels of their endpoint vertices. Mixed-colour edges report vesicles of different origin that lie within the cut-off distance, providing an automated measure of vesicles that have entered a new cell or fused with vesicles from another cell. For each image, the software computed the total number of same-colour and mixed-colour edges, vertex-level mixedness statistics (fraction of vesicles participating in mixed edges) and histograms of mixed-edge counts per vesicle, and summarized these metrics across images and biological replicates to quantify TM184C-dependent vesicle transfer. All accompanying code is available on GitHub.Statistics and reproducibilityStatistical analyses were performed in GraphPad Prism v.10.4.0. β-Arrestin and G-protein recruitment data (Fig. 2a–c,e–f,h–i) were compared using an unpaired, two-tailed Welch’s t-test; bpp score comparisons (Fig. 4e–g; 28–60 ROIs per condition, pooled from two independent biological replicates, with exact per condition n in each legend and the Source Data) used one-way and two-way ANOVA as indicated in the corresponding legend; outliers were first removed using the ROUT method (Q = 1%) in GraphPad Prism. Quantitative data are presented as mean ± s.d. (heatmap panels show the mean), with the exact n, test and P value reported in each figure legend. No statistical tests were applied to experiments with n < 3; the phosphosite mass spectrometry analysis (Fig. 2j) was performed once as a discovery screen.Immunoblots and micrographs are representative of independent experiments with similar results: Fig. 2g,k,l, n = 3; Fig. 4a–d, n > 20; Fig. 5a–e, n = 6. For all other representative images, the number of independent experiments (n = 3) is stated in the corresponding figure legend.No statistical method was used to predetermine sample size. No data were excluded from analyses. Experiments were not randomized and investigators were not blinded to allocation during experiments or outcome assessment.Ethics statementThe patient-derived GBM cells and the L1 patient-derived xenograft line were obtained as de-identified, established lines from the Ivan and Bayik laboratories, respectively, where they were derived from human tumour specimens under those institutions’ Institutional Review Board-approved protocols with informed consent. No human material was collected for the present study.Reporting summaryFurther information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
TM184C is a GPCR-like regulator of intercellular exchange and autophagy - Nature
TM184C—an ancient G-protein-coupled receptor-like superdark protein involved in regulation of autophagy, intercellular connectivity and material exchange—underscores the promise of exploring the understudied human proteome and beyond.






