# Gene ID Formatting to match Reference

**URL:** <https://community.brain-map.org/t/gene-id-formatting-to-match-reference/3730>\
**Category:** MapMyCells\
**Created:** [October 4, 2024, 5:17pm UTC](https://community.brain-map.org/t/gene-id-formatting-to-match-reference/3730 "2024-10-04T17:17:49Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![mmond](https://avatars.discourse-cdn.com/v4/letter/m/3bc359/32.png) [@mmond](https://community.brain-map.org/u/mmond)\
**Post date:** [October 4, 2024, 5:17pm UTC](https://community.brain-map.org/t/gene-id-formatting-to-match-reference/3730/1 "2024-10-04T17:17:49Z")

</div>

Hi Everyone,

I’m having an error getting my genes to map to those in the Mouse whole brain reference. I get this error: “RuntimeError: After comparing query data to reference data, no valid marker genes could be found at any level in the taxonomy.”

I’ve confirmed that the reference JSON uses Ensembl IDs and those are contained in my H5AD. However, I think my issue is arising from the way they are saved under gene\_ids. The Ensembl IDs are present, but the mapper is only referencing the first column. I feel like its an easy fix but I’m stumped, any guidance?

![image](https://canada1.discourse-cdn.com/flex027/uploads/brainobservatory/original/2X/2/21904f083ed3e826eb717dc37da72b5af2a1fedf.png)

---

<div class="post-metadata">

**Author:** ![danielsf](https://yyz1.discourse-cdn.com/flex027/user_avatar/community.brain-map.org/danielsf/32/209_2.png) [@danielsf](https://community.brain-map.org/u/danielsf)\
**Post date:** [October 4, 2024, 5:55pm UTC](https://community.brain-map.org/t/gene-id-formatting-to-match-reference/3730/2 "2024-10-04T17:55:15Z")

</div>

Hello @mmond,

The MapMyCells code looks for gene identifiers in the `index` of the `var` dataframe. Right now, it looks like your dataframe is such that

```auto
var.index.values = ['Xkr4', 'Gm1992', 'Gm19938'...]

```

and you need to be in a state where

```auto
var.index.values = ['ENSMUSG00000051951', 'ENSMUSG00000089699', ...]

```

Off-the-cuff, the easiest way to make this transformation would be

```auto
import anndata
src = anndata.read_h5ad('/path/to/original_file.h5ad')

src_var = src.var
dst_var = src_var.reset_index().set_index('gene_ids')

dst = anndata.AnnData(X=src.X, obs=src.obs, var=dst_var)
dst.write_h5ad('/path/to/reformatted_file.h5ad')

```

I am a little confused that you are having this problem. The online MapMyCells tool has a step that infers ENSEMBL IDs from gene symbols, in the event that `var` is indexed on gene symbols. Are you using the online app, or running the code locally\*? If you are running the online MapMyCells tool, would you mind posting the full run-ID when/if you encounter this error again. I would be fascinated to see what it is not inferring ENSEMBL IDs from your gene symbols.

\*if you are running the code locally, I am not surprised. The step that transforms gene symbols to Ensembl IDs is in a separate data validator module that isn’t a default part of the pipeline when running the code locally.

---

<div class="post-metadata">

**Author:** ![mmond](https://avatars.discourse-cdn.com/v4/letter/m/3bc359/32.png) [@mmond](https://community.brain-map.org/u/mmond)\
**Post date:** [October 4, 2024, 6:19pm UTC](https://community.brain-map.org/t/gene-id-formatting-to-match-reference/3730/3 "2024-10-04T18:19:11Z")

</div>

That seems to have done it. Running now, thanks. I am running locally, so that must be the issue.
