# Error to uploading my H5ad file on Mapmycell

**URL:** <https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350>\
**Category:** MapMyCells\
**Created:** [May 28, 2024, 2:36pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350 "2024-05-28T14:36:12Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![hansol](https://avatars.discourse-cdn.com/v4/letter/h/958977/32.png) [@hansol](https://community.brain-map.org/u/hansol)\
**Post date:** [May 28, 2024, 2:36pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/1 "2024-05-28T14:36:12Z")

</div>

dear community  
I tried to use Mapmycell  
here is the error

1. I have RDS file after seurat (see how my file looks like, screenshot\_)  

2. so I transformed it to h5ad with below commend:  
SaveH5Seurat(seu\_scratch\_p56, filename = “seu\_scratch\_p56.h5Seurat”)  
Convert(“seu\_scratch\_p56.h5Seurat”, dest = “h5ad”)

3. now I uploaded this h5ad file on Mapmycell

4. I got error : Mapping Failed

###### Use log files for troubleshooting MapMyCells issues. Post them [in the community forums](https://community.brain-map.org/c/how-to/mapmycells/20) for further assistance.

###### Mapping algorithm failed because of application errors.

###### Please confirm that your input data is in cell (rows) by gene (columns) format.

###### Run ID: 1716906071516-d79e81b5-b84d-4e21-a70f-4b49c71bfcbb

###### Post in the [community forum](https://community.brain-map.org/c/how-to/mapmycells/20) for help

may I ask your help for me to fix the error?  
thanks

---

<div class="post-metadata">

**Author:** ![danielsf](https://yyz1.discourse-cdn.com/flex027/user_avatar/community.brain-map.org/danielsf/32/209_2.png) [@danielsf](https://community.brain-map.org/u/danielsf)\
**Post date:** [May 28, 2024, 3:45pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/2 "2024-05-28T15:45:25Z")

</div>

Hi @hansol,

It looks like the full text of your error message is

> The ‘X’ field in this h5ad file lacks the ‘encoding-type’ metadata field. That field is necessary for this software to determine if this data is stored as a sparse or dense matrix. Please see the anndata specification here [On-disk format — anndata 0.1.dev50+g77581cc documentation](https://anndata.readthedocs.io/en/latest/fileformat-prose.html)

This means that your H5AD file is missing some metadata that is a part of the AnnData specification. I suspect that, for whatever reason, the `Convert` function you are invoking is not writing that metadata out to the output h5ad file.

I’ve consulted with some R users on our team. They recommend you write out the H5AD file directly following [the example here](https://portal.brain-map.org/explore/file-requirements-and-limits). Specifically, you will need to do something like

```auto
library(anndata) # Load library
# (First, extra count matrix from Seurat object and ensure it has row and column names.)
genes <- colnames(counts) 
samples <- rownames(counts)
sparse_counts <- as(counts, "dgCMatrix") # Convert to sparse matrix, if not already
countAD <- AnnData(X = sparse_counts, # Create the anndata object
                   var = data.frame(genes=genes,row.names=genes),
                   obs = data.frame(samples=samples,row.names=samples))
write_h5ad(countAD, "counts.h5ad", compression='gzip') # Write it out as h5ad

```

Does this help?

---

<div class="post-metadata">

**Author:** ![hansol](https://avatars.discourse-cdn.com/v4/letter/h/958977/32.png) [@hansol](https://community.brain-map.org/u/hansol)\
**Post date:** [May 28, 2024, 4:08pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/3 "2024-05-28T16:08:44Z")

</div>

thanks I managed almost but to run ANNDATA

python311.dll module error (python311.dll - The specified module could not be found.) happened  
i am wondering as I am using R not Python… but why it caused this?  
thanks a lot

---

<div class="post-metadata">

**Author:** ![danielsf](https://yyz1.discourse-cdn.com/flex027/user_avatar/community.brain-map.org/danielsf/32/209_2.png) [@danielsf](https://community.brain-map.org/u/danielsf)\
**Post date:** [May 28, 2024, 5:29pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/4 "2024-05-28T17:29:00Z")

</div>

`anndata` is a python library. Running it in R means using the R `reticulate` library to call python from R so that you have access to `anndata`’s functionality. The error message you quoted makes it sound (to me) like python did not get installed on your system.

How did you install `anndata` in your R environment?

[Here is](https://cran.r-project.org/web/packages/anndata/readme/README.html) the documentation I usually consult on running `anndata` from R. It gives a few pointers about making sure everything is installed correctly. Is this what you followed (more or less)?

---

<div class="post-metadata">

**Author:** ![hansol](https://avatars.discourse-cdn.com/v4/letter/h/958977/32.png) [@hansol](https://community.brain-map.org/u/hansol)\
**Post date:** [May 28, 2024, 5:59pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/5 "2024-05-28T17:59:07Z")

</div>

thanks for quick reply

Yes I have Python on my system. Also I used the link you sent me to re-install Anndata

You can install `anndata` for R from CRAN as follows:

```auto
install.packages("anndata")

```

Normally, reticulate should take care of installing Miniconda and the Python anndata.

If not, try running:

```auto
reticulate::install_miniconda()
anndata::install_anndata()

```

still same error… do you have any other way for me to change CSV file to H5ad file to upload to your mapmycell? I have gene in column/ cell in row in CSV form

---

<div class="post-metadata">

**Author:** ![hansol](https://avatars.discourse-cdn.com/v4/letter/h/958977/32.png) [@hansol](https://community.brain-map.org/u/hansol)\
**Post date:** [May 28, 2024, 6:38pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/6 "2024-05-28T18:38:24Z")

</div>

> countAD ← AnnData(X = sparse\_counts, # Create the anndata object

- 

```
               var = data.frame(genes=genes,row.names=genes),

```

- 

```
               obs = data.frame(samples=samples,row.names=samples))

```

Error in py\_get\_attr\_impl(x, name, silent) :  
AttributeError: ‘module’ object has no attribute ‘remap\_output\_streams’  
In addition: Warning message:  
In py\_initialize(config$python, config$libpython, config$pythonhome, :  
Python 2 reached EOL on January 1, 2020. Python 2 compatability be removed in an upcoming reticulate release.

additionally… when I used another PC

---

<div class="post-metadata">

**Author:** ![hansol](https://avatars.discourse-cdn.com/v4/letter/h/958977/32.png) [@hansol](https://community.brain-map.org/u/hansol)\
**Post date:** [May 28, 2024, 7:03pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/7 "2024-05-28T19:03:21Z")

</div>

Also when I used Python (Jupyter note)  
its not working… I think something i am stuck…  
converting CSV file to H5ad, i dont know why it is so hard…

 ![image](https://canada1.discourse-cdn.com/flex027/uploads/brainobservatory/original/2X/d/db169e5a8b70bd625da20681e9cc6f4179d9df33.png)

---

<div class="post-metadata">

**Author:** ![danielsf](https://yyz1.discourse-cdn.com/flex027/user_avatar/community.brain-map.org/danielsf/32/209_2.png) [@danielsf](https://community.brain-map.org/u/danielsf)\
**Post date:** [May 28, 2024, 7:37pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/8 "2024-05-28T19:37:19Z")

</div>

Sorry. I did not realize you were comfortable using Python.

Can you regenerate your Jupyter notebook error, but grab a screenshot that includes the entire pink window? You clipped off the important part of the error message, unfortunately.

Thanks.

---

<div class="post-metadata">

**Author:** ![hansol](https://avatars.discourse-cdn.com/v4/letter/h/958977/32.png) [@hansol](https://community.brain-map.org/u/hansol)\
**Post date:** [May 28, 2024, 7:42pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/9 "2024-05-28T19:42:59Z")

</div>

![image](https://canada1.discourse-cdn.com/flex027/uploads/brainobservatory/original/2X/b/b40ed56743ac7bc3a8127de621a52e2d696fafa4.png)

---

<div class="post-metadata">

**Author:** ![danielsf](https://yyz1.discourse-cdn.com/flex027/user_avatar/community.brain-map.org/danielsf/32/209_2.png) [@danielsf](https://community.brain-map.org/u/danielsf)\
**Post date:** [May 28, 2024, 7:43pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/10 "2024-05-28T19:43:08Z")

</div>

Though, if you don’t want to go to that trouble, there is a `read_csv` function in the `anndata` library.

If, for instance, your data is in a file called `junk.txt` that looks like

```auto
,g1,g2,g3
c1,1.0,2.0,3.0
c2,4.0,5.0,6.0

```

you could run

```auto
>>> import anndata
>>> a = anndata.read_csv('junk.txt', first_column_names=True)
>>> a.write_h5ad('junk.h5ad')
>>> 

```

---

<div class="post-metadata">

**Author:** ![hansol](https://avatars.discourse-cdn.com/v4/letter/h/958977/32.png) [@hansol](https://community.brain-map.org/u/hansol)\
**Post date:** [May 28, 2024, 7:46pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/11 "2024-05-28T19:46:26Z")

</div>

![image](https://canada1.discourse-cdn.com/flex027/uploads/brainobservatory/original/2X/7/748db33911990ee65366a79ad70d2419fb5c2454.png)

---

<div class="post-metadata">

**Author:** ![danielsf](https://yyz1.discourse-cdn.com/flex027/user_avatar/community.brain-map.org/danielsf/32/209_2.png) [@danielsf](https://community.brain-map.org/u/danielsf)\
**Post date:** [May 28, 2024, 7:52pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/12 "2024-05-28T19:52:01Z")

</div>

the problem with your Jupyter notebook is here

 ![Screenshot 2024-05-28 at 12.43.33 PM](https://canada1.discourse-cdn.com/flex027/uploads/brainobservatory/original/2X/2/2699568b9a668d8ee845764363b45df71a5a46ee.png)

The cell-by-gene array which you store as `result` has strings in the first column. It should just be the numerical values. The `NaN` values in that matrix aren’t ideal, either. Is this log-normalized data?

I’m also concerned that you created your `anndata.AnnData` object here

 ![Screenshot 2024-05-28 at 12.50.06 PM](https://canada1.discourse-cdn.com/flex027/uploads/brainobservatory/original/2X/7/7cb962c4901f7f7b23484d6358cafdcafe00280b.png)

without specifying the `obs` or `var` dataframes. `obs` is a dataframe that identifies each cell (row) in your cell-by-gene matrix. `var` is a dataframe that identifies each gene (column) in your cell-by-gene matrix. It may be illuminating to look at the second “in Python” examle on [this page](https://portal.brain-map.org/explore/file-requirements-and-limits) (I know I already pointed you to this page, but the first example starts from a more complicated data model than I think you are working with). Here is the part of that page that does the simplest creation of `obs` and `var`.

 ![Screenshot 2024-05-28 at 12.50.39 PM](https://canada1.discourse-cdn.com/flex027/uploads/brainobservatory/original/2X/a/a31f759a36657894c4ee52d30e70494c40a83278.png)

The creation of the `AnnData` object then looks like

 ![Screenshot 2024-05-28 at 12.51.26 PM](https://canada1.discourse-cdn.com/flex027/uploads/brainobservatory/original/2X/1/1bdd5a1158c68df2f6523a2b5665e26f7abbf013.png)

---

<div class="post-metadata">

**Author:** ![hansol](https://avatars.discourse-cdn.com/v4/letter/h/958977/32.png) [@hansol](https://community.brain-map.org/u/hansol)\
**Post date:** [May 28, 2024, 8:09pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/13 "2024-05-28T20:09:42Z")

</div>

> [@danielsf](#):
>
> ```auto
> >>> import anndata
> >>> a = anndata.read_csv('junk.txt', first_column_names=True)
> >>> a.write_h5ad('junk.h5ad')
> 
> ```

it worked i will try to upload now! thanks

---

<div class="post-metadata">

**Author:** ![hansol](https://avatars.discourse-cdn.com/v4/letter/h/958977/32.png) [@hansol](https://community.brain-map.org/u/hansol)\
**Post date:** [May 28, 2024, 8:57pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/14 "2024-05-28T20:57:41Z")

</div>

this worked perfectly  
so now I made

comparison = table(mapping$subclass\_name,seu\_scratch\_p56@meta.data[[“seurat\_clusters”]])  
heatmap(comparison, ylab = “Mapping”, xlab = “Clustering”,margins = c(10,10))

 ![image](https://canada1.discourse-cdn.com/flex027/uploads/brainobservatory/original/2X/3/3dba8d4c197dadaeddbfdf615fb8e1df1d02da27.png)

and then now

dataSeurat ← CreateSeuratObject(counts = seu\_scratch\_p56@assays[[“RNA”]]@counts, meta.data = mapping)

library( Seurat)

# Standard Seurat pipeline

dataSeurat ← NormalizeData(dataSeurat, verbose = FALSE)  
dataSeurat ← FindVariableFeatures(dataSeurat, verbose = FALSE)  
dataSeurat ← ScaleData(dataSeurat, verbose = FALSE)  
dataSeurat ← RunPCA(dataSeurat, verbose = FALSE)  
dataSeurat ← RunUMAP(dataSeurat, dims = 1:10, verbose = FALSE)

DimPlot(dataSeurat, reduction = “umap”, group.by=“subclass\_name”, label=TRUE) + NoLegend()  
I tried this and it showed the error :

Error in `[.data.frame`(data, , group) : undefined columns selected  
In addition: Warning message:  
The following requested variables were not found: subclass\_name

may I ask what did I wrongly? thanks a lot

---

<div class="post-metadata">

**Author:** ![danielsf](https://yyz1.discourse-cdn.com/flex027/user_avatar/community.brain-map.org/danielsf/32/209_2.png) [@danielsf](https://community.brain-map.org/u/danielsf)\
**Post date:** [May 28, 2024, 9:20pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/15 "2024-05-28T21:20:32Z")

</div>

This is just a guess, but there are 4-5 lines at the top of the file that you downloaded which are just metadata. These are marked with a `#` at the front, so you will probably need to tell Seurat to ignore lines that start with `#` (or open the CSV and delete those lines).

Did that work?

---

<div class="post-metadata">

**Author:** ![hansol](https://avatars.discourse-cdn.com/v4/letter/h/958977/32.png) [@hansol](https://community.brain-map.org/u/hansol)\
**Post date:** [May 29, 2024, 7:58am UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/16 "2024-05-29T07:58:36Z")

</div>

> [@hansol](#):
>
> Error in `[.data.frame`(data, , group) : undefined columns selected  
> In addition: Warning message:  
> The following requested variables were not found: subclass\_name

hi Deanel, it is not about # command uncommand things… I think

in my dataseurat I do not have subclass\_name

so after mapping  
I have mapping data from MAP MYCELL  
and as your tutorial ([MapMyCells Use Case: Single Nucleus RNA-seq from Human MTG - brain-map.org](https://portal.brain-map.org/atlases-and-data/bkp/mapmycells/mapmycells-use-case-single-nucleus-rnaseq-from-human-mtg)) suggested I go through  
especially, → #7 To visualize the mapping results, we need both the mapping results and the original query cellxgene matrix for comparison. If this is not already read in, you can read it in from the anndata object uploaded to MapMyCells.

```auto
### Since the query data corresponds to dataQC above, we will call it dataQC again
dataQC_h5ad <- read_h5ad('Hodge2019.h5ad')
dataQC <- t(as.matrix(dataQC_h5ad$X))
rownames(dataQC) <- rownames(dataQC_h5ad$var)
colnames(dataQC) <- rownames(dataQC_h5ad$obs)

```

here as my R did not work with anndata anyway i just used my seuratObj (original one)

```auto
#Create the Seurat object
dataSeurat <- CreateSeuratObject(counts = dataQC, meta.data = mapping)

```

here instead of dataQC I used seu\_scratch\_p56@assays[[“RNA”]]@counts  
seu\_scratch\_p56 (myseurat obj) which I used to create anndata with below command:

```auto
gene_matrix <- seu_scratch_p56[["RNA"]]$data
write.csv(gene_matrix, file = "gene_matrix.csv", row.names = TRUE)

```

thanks for helping me!

---

<div class="post-metadata">

**Author:** ![hansol](https://avatars.discourse-cdn.com/v4/letter/h/958977/32.png) [@hansol](https://community.brain-map.org/u/hansol)\
**Post date:** [May 29, 2024, 1:33pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/17 "2024-05-29T13:33:13Z")

</div>

> [@hansol](#):
>
> `dataQC <- t(as.matrix(dataQC_h5ad$X))`

Instead of Anndata format, what I can extract directly from Seurat Object in order to replace here dataQC\_h5ad$X (what is it inside)? thanks!

---

<div class="post-metadata">

**Author:** ![danielsf](https://yyz1.discourse-cdn.com/flex027/user_avatar/community.brain-map.org/danielsf/32/209_2.png) [@danielsf](https://community.brain-map.org/u/danielsf)\
**Post date:** [May 29, 2024, 3:11pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/18 "2024-05-29T15:11:12Z")

</div>

I’m an exclusive python user, so I don’t have any specific knowledge of R or Seurat. My understanding is that this line

```auto
dataSeurat <- CreateSeuratObject(counts = dataQC, meta.data = mapping)

```

should have joined the original cell-by-gene data with the mapping from MapMyCells. What columns _are_ in your `dataSeurat` object?

---

<div class="post-metadata">

**Author:** ![hansol](https://avatars.discourse-cdn.com/v4/letter/h/958977/32.png) [@hansol](https://community.brain-map.org/u/hansol)\
**Post date:** [May 29, 2024, 3:34pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/19 "2024-05-29T15:34:06Z")

</div>

I dont know whether it is matter of language…and also tutorial is written based on R …

then could you let me know

1. what is content of this dataQC (what is in X)? dataQC ← t(as.matrix(dataQC\_h5ad$X))

2. is there any better tutorial or vignette for Python user? After mapping how they can compare the mapping with Allen ?

thanks a lot

 ![image](https://canada1.discourse-cdn.com/flex027/uploads/brainobservatory/original/2X/2/20d98cf7eef9e5e0ca826c31caf9cf36019e3b1d.png)

---

<div class="post-metadata">

**Author:** ![danielsf](https://yyz1.discourse-cdn.com/flex027/user_avatar/community.brain-map.org/danielsf/32/209_2.png) [@danielsf](https://community.brain-map.org/u/danielsf)\
**Post date:** [May 29, 2024, 4:32pm UTC](https://community.brain-map.org/t/error-to-uploading-my-h5ad-file-on-mapmycell/3350/20 "2024-05-29T16:32:17Z")

</div>

The `X` matrix is a convention in h5ad files. Specifically, in an h5ad file

- `X` is the cell-by-gene expression matrix (just an array of floats or integers where each row is a cell and each column is a gene)
- `obs` is a dataframe containing the metadata for each cell. Each row in this dataframe corresponds to a row in `X`.
- `var` is a dataframe containing the metadata for each gene. Eachrow in this dataframe corresponds to a column in `X`.

I’m not sure we have any visualization tutorials written in python, for better or worse.

MapMyCells should have given you a CSV file in which each row is a cell in your dataset and each column is either a cell type assignment or a quality metric for a cell type assignment (see documentation [here](https://github.com/AllenInstitute/cell_type_mapper/blob/main/docs/output.md#csv-output-file)). I’d recommend comparing that CSV file to your original dataset using whatever data analysis tools you are most comfortable with.
