Bioinformatics

By analyzing extensive curated metabolomics, proteomics, and transcriptomics data related to creatine metabolism, this section synthesizes key insights into creatine’s functions at the cellular and tissue levels, as well as its potential implications in health and disease. The Bioinformatics section of CREAS employs computational tools to investigate the molecular mechanisms of creatine.

1. Metabolomics

1.1. Creatine

1.1.1. Description

Creatine is a naturally occurring compound. It belongs to the class of organic compounds known as alpha amino acids and derivatives. These are amino acids in which the amino group is attached to the carbon atom immediately adjacent to the carboxylate group (alpha carbon), or a derivative thereof. Creatine is found in all vertebrates where it facilitates recycling of adenosine triphosphate (ATP). Its primary metabolic role is to combine with a phosphoryl group, via the enzyme creatine kinase, to generate phosphocreatine, which is used to regenerate ATP. Most of the human body’s total creatine and phosphocreatine stores are found in skeletal muscle (95%), while the remainder is distributed in the blood, brain, testes, and other tissues.

1.1.2. Metabolite Identification
    • Chemical formula: C4H9N3O2
    • Monoisotopic Molecular Weight: 131.069476547
    • IUPAC name: 2-(N-methylcarbamimidamido)acetic acid
    • CAS Registry Number: 57-00-1
    • SMILES: CN(CC(O)=O)C(N)=N

1.1.3. NMR Structure of Creatine
1.1.4. Spectra
  • GC-MS spectra

    Spectrum type: Experimental GC-MS; Description; GC-MS Spectrum – Creatine GC-EI-TOF (Non-derivatized); Splash key: splash10-0002-0900000000-f89f340e776ae37776b3; Deposition date: 2018-05-18; Source: HMDB team, MONA, MassBank
  • MS-MS spectra

    Spectrum type: Experimental LC-MS/MS; Description; LC-MS/MS Spectrum – Creatine Quattro_QQQ 40V, Positive-QTOF (Annotated); Splash key: splash10-0006-9000000000-13b7d28e8f5be1c5dada; Deposition date: 2012-07-24; Source: HMDB team, MONA
  • NMR spectra

    Spectrum type: Experimental 1D NMR; Description; 1H NMR Spectrum (1D, 500 MHz, H2O, experimental); Deposition date: 2012-12-04; Source: Wishart Lab
1.1.5. Biological Properties
  • Cellular locations: Cytoplasm, Extracellular and Mitochondria
  • Biospecimen Locations: Blood, Breast Milk, Cerebrospinal Fluid (CSF), Feces, Saliva, Sweat, Urine.
  • Tissue locations: Adipose Tissue, Bladder, Brain, Epidermis, Fibroblasts, Heart, Intestine, Kidney, Neuron, Placenta, Platelet, Prostate, Skeletal Muscle, Spleen, Testis
1.1.6. Enzymes
  • Glycine amidinotransferase
    • General function: Amino acid transport and metabolism
    • Gene Name: GATM
    • Uniprot ID: P50440
    • Molecular weight: 48455.01
  • Guanidinoacetate N-methyltransferase
    • General function: Involved in guanidinoacetate N-methyltransferase activity
    • Gene Name: GAMT
    • Uniprot ID: Q14353
    • Molecular weight: 26317.925
  • Creatine kinase S-type, mitochondrial
    • General function: Involved in kinase activity
    • Gene Name: CKMT2
    • Uniprot ID: P17540
    • Molecular weight: 47504.08
  • Creatine kinase U-type, mitochondrial
    • General function: Involved in kinase activity
    • Gene Name: CKMT1A
    • Uniprot ID: P12532
    • Molecular weight: 47036.3
  • Creatine kinase B-type
    • General function: Involved in kinase activity
    • Gene Name: CKB
    • Uniprot ID: P12277
    • Molecular weight: 42643.95
  • Creatine kinase M-type
    • General function: Involved in kinase activity
    • Gene Name: CKM
    • Uniprot ID: P06732
    • Molecular weight: 43100.91
  • Solute carrier family 6 member 8
    • General function: Involved in neurotransmitter:sodium symporter activity
    • Gene Name: SLC6A8
    • Uniprot ID: P48029
    • Molecular weight: 69483.91
  • Solute carrier family 22 member 5
    • General function: Involved in ion transmembrane transporter activity. A creatine efflux transporter in oligodendrocytes
    • Gene Name: SLC22A5
    • Uniprot ID: O76082

1.2. Phosphocreatine

1.2.1. Description

Phosphocreatine, also known as phosphorylcreatine or creatine phosphate (PCr), is a phosphorylated creatine molecule that serves as a rapidly mobilizable reserve of high-energy phosphates in skeletal muscle, myocardium and the brain to recycle adenosine triphosphate, the energy currency of the cell. Phosphocreatine undergoes irreversible cyclization and dehydration to form creatinine at a fractional rate of 0.026 per day, thus forming approximately 2 g creatinine/day in an adult male. This is the amount of creatine that must be provided either from dietary sources or by endogenous synthesis to maintain the body pool of (creatine and) phosphocreatine.

1.2.2. Metabolite Identification
    • Chemical formula: C4H10N3O5P
    • Monoisotopic Molecular Weight: 211.035806957
    • IUPAC name: 2-(N-methyl-N’-phosphonocarbamimidamido)acetic acid
    • CAS Registry Number: 67-07-2
    • SMILES: CN(CC(O)=O)C(=N)NP(O)(O)=O
1.2.3. NMR Structure of Phosphocreatine

1.2.4. Spectra
  • MS-MS spectra

    Spectrum type: Experimental LC-MS/MS; Description; LC-MS/MS Spectrum – Phosphocreatine Quattro_QQQ 40V, Positive-QTOF (Annotated); Splash key: splash10-004i-9000000000-0b9440fbeacd93f8e2c6; Deposition date: 2012-07-24; Source: HMDB team, MONA
  • NMR spectra

    Spectrum type: Experimental 1D NMR; Description; 1H NMR Spectrum (1D, 500 MHz, H2O, experimental); Deposition date: 2012-12-04; Source: Wishart Lab
1.2.5. Biological Properties

• Cellular locations: Cytoplasm, Extracellular and Mitochondria
• Biospecimen Locations: Blood, Breast Milk, Urine.
• Tissue locations: Adipose Tissue, Basal Ganglia, Brain, Epidermis, Fibroblasts, Kidney, Neuron, Placenta, Skeletal Muscle, Testis

1.2.6. Enzymes
  • Creatine kinase S-type, mitochondrial
    • General function: Involved in kinase activity
    • Gene Name: CKMT2
    • Uniprot ID: P17540
    • Molecular weight: 47504.08
  • Creatine kinase U-type, mitochondrial
    • General function: Involved in kinase activity
    • Gene Name: CKMT1A
    • Uniprot ID: P12532
    • Molecular weight: 47036.3
  • Creatine kinase B-type
    • General function: Involved in kinase activity
    • Gene Name: CKB
    • Uniprot ID: P12277
    • Molecular weight: 42643.95
  • Creatine kinase M-type
    • General function: Involved in kinase activity
    • Gene Name: CKM
    • Uniprot ID: P06732
    • Molecular weight: 43100.91

1.3. Creatinine

1.3.1. Description

Creatinine or creatine anhydride is a member of the class of compounds known as imidazolidinones. Imidazolidinones are a class of 5-membered ring heterocycles structurally related to imidazole. Creatinine can also be classified as an amino acid derivative.  Creatinine arises from the production of creatine. In particular, the loss of a water molecule from creatine results in the formation of creatinine. Creatinine is transferred to the kidneys by blood plasma, whereupon it is eliminated from the body by glomerular filtration and partial tubular excretion. Creatinine is usually produced at a fairly constant rate by the body, which is roughly proportional to muscle mass and body size

1.3.2. Metabolite Identification
    • Chemical formula: C4H7N3O
    • Monoisotopic Molecular Weight: 113.058911861
    • IUPAC name: 2-imino-1-methylimidazolidin-4-one
    • CAS Registry Number: 60-27-5
    • SMILES: CN1CC(=O)NC1=N
1.3.3. NMR Structure of Phosphocreatine

1.3.4. Spectra
  • MS-MS spectra

    Spectrum type: Experimental GC-MS; Description; GC-MS Spectrum – Creatinine GC-MS (Non-derivatized); Splash key: splash10-014i-0901000000-bd882951c92733ecbc9e; Deposition date: 2017-09-12; Source: HMDB team, MONA, MassBank
  • NMR spectra

    Spectrum type: Experimental LC-MS/MS Spectrum – Creatinine Quattro_QQQ 40V, Positive-QTOF (Annotated); Splash key: splash10-0006-9000000000-bfa35af1a437cbc8b5ec; Deposition date: 2012-07-24; Source: HMDB team, MONA
1.3.5. Biological Properties
  • Cellular locations: Cytoplasm
  • Biospecimen Locations: Amniotic Fluid, Blood, Breast Milk, Cerebrospinal Fluid (CSF), Feces, Saliva, Sweat, Urine
  • Tissue locations: Adipose Tissue, Bladder, Fibroblasts, Kidney, Liver, Neuron, Pancreas, Placenta, Platelet, Prostate, Skeletal Muscle, Spleen, Testis, Thyroid Gland

Data extracted from the Human Metabolome Database (HMDB) version 5.0, a freely available electronic database containing detailed information about small molecule metabolites found in the human body. HMDB is a project is supported by the Canadian Institutes of Health Research, Canada Foundation for Innovation, and by The Metabolomics Innovation Centre (TMIC), a nationally-funded research and core facility that supports a wide range of cutting-edge metabolomic studies.


References:

Wishart DS, Tzur D, Knox C, et al., HMDB: the Human Metabolome Database. Nucleic Acids Res. 2007 Jan;35(Database issue):D521-6. 17202168

Wishart DS, Knox C, Guo AC, et al., HMDB: a knowledgebase for the human metabolome. Nucleic Acids Res. 2009 37(Database issue):D603-610. 18953024

Wishart DS, Jewison T, Guo AC, Wilson M, Knox C, et al., HMDB 3.0 — The Human Metabolome Database in 2013. Nucleic Acids Res. 2013. Jan 1;41(D1):D801-7. 23161693

Wishart DS, Feunang YD, Marcu A, Guo AC, Liang K, et al., HMDB 4.0 — The Human Metabolome Database for 2018. Nucleic Acids Res. 2018. Jan 4;46(D1):D608-17. 29140435

Wishart DS, Guo AC, Oler E, et al., HMDB 5.0: the Human Metabolome Database for 2022. Nucleic Acids Res. 2022. Jan 7;50(D1):D622–31. 34986597

2. RNA and Protein Tissue Expression

Below is an overview of RNA and protein expression data generated in the Human Protein Atlas project. Analyzed tissues are divided into color-coded groups according to which functional features they have in common. For each group, a list of included tissues is accessed by clicking on the group name, group symbol, RNA bar, or protein bar. Subsequent selection of a particular tissue in this list links to the image data page.

The immunohistochemistry image panel to the right display selected tissues that give a visual summary of the protein expression profile. To the left, the two human bodies provide an anatomical display of the expression levels and detection of mRNA in the analyzed organs. Data taken from The Human Protein Atlas – no changes were made. Licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (https://creativecommons.org/licenses/by-sa/4.0/).

2.1. GATM
  • Description: Glycine amidinotransferase
  • Tissue specificity (RNA): Group enriched (Kidney, Liver, Pancreas)
  • Tau specificity score (RNA):  0.55
  • Protein evidence: Evidence at protein level
  • Protein expression: Granular cytoplasmic expression in several different tissue types, mainly in kidney, liver and pancreas.
  • Data reliability description: Medium consistency between antibody staining and RNA expression data.
  • Reliability score: Enhanced
  • Antibodies: HPA026077

2.2. GAMT
  • Description: Guanidinoacetate N-methyltransferase
  • Tissue specificity (RNA): Tissue enhanced (Liver, Skeletal muscle, Tongue)
  • Tau specificity score (RNA):  0.52
  • Protein evidence: Evidence at protein level
  • Protein expression:    Cytoplasmic and membranous expression in several tissues.
  • Data reliability description: Low consistency between antibody staining and RNA expression data.
  • Reliability score: Approved
  • Antibodies: HPA051806

2.3. CKB
  • Description: Creatine kinase B
  • Tissue specificity (RNA): Tissue enhanced (Brain)
  • Tau specificity score (RNA):  0.36
  • Protein evidence: Evidence at protein level
  • Protein expression: Cytoplasmic expression in parietal cells in stomach, and cells in the CNS.
  • Data reliability description: High consistency between antibody staining and RNA expression data.
  • Reliability score: Enhanced
  • Antibodies: HPA001254 , CAB047313

2.4. CKM
  • Description: Creatine kinase, M-type
  • Tissue specificity (RNA): Group enriched (Skeletal muscle, Tongue)
  • Tau specificity score (RNA):  0.81
  • Protein evidence: Evidence at protein level
  • Protein expression: Cytoplasmic expression in skeletal muscle and heart.
  • Data reliability description: High consistency between antibody staining and RNA expression data.
  • Reliability score: Enhanced
  • Antibodies: HPA047859

2.5. CKMT1A
  • Description: Creatine kinase, mitochondrial 1A
  • Tissue specificity (RNA): Tissue enhanced (Brain, Esophagus, Intestine)
  • Tau specificity score (RNA):  0.59
  • Protein evidence: Evidence at protein level
  • Protein expression: Granular cytoplasmic expression in several different tissue types, mainly in gastrointestinal tract and squamous epithelium.
  • Data reliability description: Low consistency between antibody staining and RNA expression data. Caution, targets protein from more than one gene. Presumed off target binding observed and disregarded.
  • Reliability score: Uncertain
  • Antibodies: HPA043491

2.6. CKMT1B
  • Description: Creatine kinase, mitochondrial 1B
  • Tissue specificity (RNA): Tissue enhanced (Brain, Esophagus, Intestine)
  • Tau specificity score (RNA):  0.62
  • Protein evidence: Evidence at protein level
  • Protein expression: Granular cytoplasmic expression in several different tissue types, mainly in gastrointestinal tract and squamous epithelium.
  • Data reliability description: Low consistency between antibody staining and RNA expression data. Caution, targets protein from more than one gene. Presumed off target binding observed and disregarded.
  • Reliability score: Uncertain
  • Antibodies: HPA043491

2.7. CKMT2
  • Description: Creatine kinase, mitochondrial 2
  • Tissue specificity (RNA): Group enriched (Heart muscle, Skeletal muscle, Tongue)
  • Tau specificity score (RNA):  0.69
  • Protein evidence: Evidence at protein level
  • Protein expression: Cytoplasmic expression mainly in heart and skeletal muscle..
  • Data reliability description: Medium consistency between antibody staining and RNA expression data. Presumed off target binding observed and disregarded.
  • Reliability score: Enhanced
  • Antibodies: HPA051880

2.8. SLC6A8
  • Description: Solute carrier family 6 member 8
  • Tissue specificity (RNA): Low tissue specificity
  • Tau specificity score (RNA):  0.33
  • Protein evidence: Detected in all
  • Reliability score: Pending normal tissue annotation.

Notes:

Tissue specificity (RNA): The RNA specificity category is based on mRNA expression levels in the consensus dataset which is calculated from the RNA expression levels in samples from HPA and GTEX. The categories include tissue enriched, group enriched, tissue enhanced, low tissue specificity and not detected.

Tau specificity score (RNA): Tau specificity score is a numerical indicator of the specificity of the gene expression across cells or tissues. The value ranges from 0 and 1, where 0 indicates identical expression across all cells/tissue types, while 1 indicates expression in a single cell/tissue type.

Protein evidence: Evidence score for genes based on UniProt protein existence (UniProt evidence); neXtProt protein existence (neXtProt evidence);and a Human Protein Atlas antibody- or RNA based score (HPA evidence). The avaliable scores are evidence at protein level, evidence at transcript level, no evidence, or not avaliable.

Protein expression: A summary of the overall protein expression pattern across the analyzed normal tissues. The summary is based on knowledge-based annotation. “Estimation of protein expression could not be performed. View primary data.” is shown for genes analyzed with a knowledge-based approach where available RNA-seq and gene/protein characterization data has been evaluated as not sufficient in combination with immunohistochemistry data to yield a reliable estimation of the protein expression profile.

Data reliability description: Standardized explanatory sentences with additional information required for full understanding of the protein expression profile, based on knowledge-based and secretome-based annotation.

Reliability score: A reliability score is manually set for all genes and indicates the level of reliability of the analyzed protein expression pattern based on available RNA-seq data, protein/gene characterization data and immunohistochemical data from one or several antibodies with non-overlapping epitopes. The reliability score is based on the 44 normal tissues analyzed, and if there is available data from more than one antibody, the staining patterns of all antibodies are taken into consideration during evaluation. The reliability score is divided into Enhanced, Supported, Approved, or Uncertain, and is displayed on both Tissue resource and Cancer resource.

Antibodies: Antibodies used for this assay.

RNA expression (nTPM): RNA expression summary shows the consensus data based on normalized expression (nTPM) values from two different sources: internally generated Human Protein Atlas (HPA) RNA-seq data and RNA-seq data from the Genotype-Tissue Expression (GTEx) project. Color-coding is based on tissue groups, each consisting of tissues with functional features in common.

Protein expression (score): Each bar represents the highest expression score found in a particular group of tissues. Protein expression scores are based on a best estimate of the “true” protein expression from a knowledge-based annotation, described more in detail under Assays & annotation (https://www.proteinatlas.org/about/assays+annotation#ihk). For genes where more than one antibody has been used, a collective score is set displaying the estimated true protein expression.

3. Enrichment Analyses

We utilized a curated list of creatine metabolism-related genes/proteins—GAMT, GATM, CKB, CKM, CKMT1B, CKMT1A, CKMT2, SLC6A8, and SLC6A10P—to perform enrichment analysis. This approach helped identify key biological pathways, processes, and associations linked to creatine metabolism, providing deeper insights into its functional roles and potential implications.

    3.1. Enrichr Knowledge Graph (Enrichr-KG)

    This combines enrichment analysis with a knowledge graph data representation to query a large collection of processed datasets made of associations between genes and many biological and biomedical terms. This is an extension of the popular gene set enrichment analysis tool Enrichr, designed to enhance the integration and visualization of enrichment results across multiple gene set libraries and domains of knowledge (available at: https://maayanlab.cloud/enrichr-kg).

    The permanent link below provides access to enrichment analysis using the classical Enrichr tool, covering diverse categories such as Transcription, Pathways, Ontologies, Diseases/Drugs, Cell Types, Misc, Legacy, and Crowd. Use it to explore and extract detailed insights across these domains:

    GENE SET ENRICHMENT ANALYSIS

    References:

    • Evangelista JE, Xie Z, Marino GB, Nguyen N, Clarke DJB, Ma’ayan A. Enrichr-KG: bridging enrichment analysis across multiple libraries. Nucleic Acids Res. 2023 May 11:gkad393. https://doi.org/10.1093/nar/gkad393
    • Chen, E. Y., Tan, C. M., Kou, Y., Duan, Q., Wang, Z., Meirelles, G. V., Clark, N. R., & Ma’ayan, A. (2013). Enrichr: interactive and collaborative HTML5 gene list enrichment analysis tool. BMC bioinformatics14, 128. https://doi.org/10.1186/1471-2105-14-128
    • Kuleshov, M. V., Jones, M. R., Rouillard, A. D., Fernandez, N. F., Duan, Q., Wang, Z., Koplev, S., Jenkins, S. L., Jagodnik, K. M., Lachmann, A., McDermott, M. G., Monteiro, C. D., Gundersen, G. W., & Ma’ayan, A. (2016). Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic acids research44(W1), W90–W97. https://doi.org/10.1093/nar/gkw377 
    • Xie, Z., Bailey, A., Kuleshov, M. V., Clarke, D. J., Evangelista, J. E., Jenkins, S. L., … & Ma’ayan, A. (2021). Gene set knowledge discovery with Enrichr. Current protocols1(3), e90. https://doi.org/10.1002/cpz1.90
    3.2. Protein-protein interaction Network

    Protein-protein interactions network of creatine metabolism. The prioritized gene/proteins identified in the manual curation were submitted to the Search Tool for the Retrieval of Interacting Genes (STRING, https://string-db.org/). The colored nodes represent the results of the Markov cluster algorithm to group proteins. The colors of interactions correspond to: known from curated databases (cyan), experimentally determined (purple); predicted interactions based on gene neighborhood (green), gene fusions (red), and gene co-occurrence (dark blue); and others, such as text-mining (yellow), co-expression (black), and protein homology (light blue).

    The network is available at https://version-12-0.string-db.org/cgi/network?networkId=bwu8z9Urdb75.

    STRING is a Core Data Resource as designated by Global Biodata Coalition and ELIXIR.

    3.3. Predicted miRNA-mRNA interactions

    The DIANA-microT 2023 webserver is a computational tool developed by the DIANA Lab that predicts and archives miRNA-mRNA interactions (available at: https://dianalab.e-ce.uth.gr/microt_webserver/#/interactions). It uses the DIANA-microT-CDS algorithm to analyze miRNA sequences and their interactions with protein-coding genes, hosting over 86 million predicted interactions across six species. It also includes predictions for viral miRNAs targeting host transcripts. This resource aids in understanding miRNA roles in gene regulation and disease, integrating data from various biological levels for comprehensive analysis.

    The gene list of creatine metabolism-related genes (GAMT, GATM, CKB, CKM, CKMT1B, CKMT1A, CKMT2, SLC6A8, and SLC6A10P) was used as input data, with an interaction score threshold of 0.7, human as species, and high miRNA confidence settings.

    Reference:

    • Tastsoglou, S., Alexiou, A., Karagkouni, D., Skoufos, G., Zacharopoulou, E., & Hatzigeorgiou, A. G. (2023). DIANA-microT 2023: including predicted targets of virally encoded miRNAs. Nucleic acids research51(W1), W148–W153. https://doi.org/10.1093/nar/gkad283
    • The “ELIXIR-GR: Managing and Analysing Life Sciences Data (MIS: 5002780)” Project is co-financed by Greece and the European Union – European Regional Development Fund