Dear Allen Brain Cell Atlas / CCF registration team,
I am currently integrating the whole-brain MERFISH dataset MERFISH-C57BL6J-638850 with the Allen Mouse CCF and an existing mouse brain connectome/ontology database.
During quality control, I noticed a systematic difference between the complete MERFISH cell metadata and the publicly available CCF registration products.
For the original MERFISH dataset,
MERFISH-C57BL6J-638850/metadata/views/cell_metadata_with_cluster_annotation.csv
I obtain:
-
3,938,808 unique cells
-
59 brain sections
This agrees with the documented size of the MERFISH dataset.
However, the CCF registration products from
MERFISH-C57BL6J-638850-CCF/20231215
contain only 3,739,961 unique cells in each of the following files:
-
reconstructed_coordinates.csv -
ccf_coordinates.csv -
views/cell_metadata_with_parcellation_annotation.csv
Thus, 198,847 MERFISH cells (5.05%) are absent from all three CCF-derived products.
I compared the cell_label identifiers directly, so this difference does not result from duplicate cells or from a downstream join performed on our side.
1. Six complete sections are absent
The following six MERFISH sections are present in the original cell metadata but completely absent from the CCF-derived files:
-
C57BL6J-638850.01— 23,389 cells -
C57BL6J-638850.02— 25,792 cells -
C57BL6J-638850.03— 30,546 cells -
C57BL6J-638850.04— 43,047 cells -
C57BL6J-638850.68— 20,040 cells -
C57BL6J-638850.69— 13,963 cells
Together these account for 156,777 cells.
2. Additional cells are missing from otherwise registered sections
Another 42,070 cells are absent from the CCF products although their respective sections are otherwise represented.
These missing cells occur in 11 sections.
Some examples are:
-
.27: 8,118 / 58,039 cells missing (13.99%) -
.42: 8,051 / 75,551 (10.66%) -
.33: 8,425 / 85,849 (9.81%) -
.48: 5,915 / 77,165 (7.67%) -
.24: 2,622 / 64,243 (4.08%) -
.32: 36 / 93,525 (0.038%)
The missing cells appear to be spatially structured rather than randomly distributed.
For example, in section .27, the 1% x-coordinate quantile of the CCF-registered cells is approximately 0.96, whereas for the 8,118 missing cells it is approximately 6.80.
The missing populations are also not obviously characterized by poor transcriptomic annotation quality. Their mean average_correlation_score values are generally in the same range as the remaining dataset, and the missing cells comprise multiple classes and hundreds of clusters.
3. Public S3 resources
I also recursively inspected the public S3 resources for this dataset.
For
metadata/MERFISH-C57BL6J-638850-CCF/20231215/
I find only:
-
ccf_coordinates.csv -
reconstructed_coordinates.csv -
views/cell_metadata_with_parcellation_annotation.csv
For
image_volumes/MERFISH-C57BL6J-638850-CCF/20230630/
I find:
-
resampled_annotation.nii.gz -
resampled_annotation_boundary.nii.gz -
resampled_average_template.nii.gz
I could not find the global affine transform, section-wise affine transforms, deformation fields, displacement fields, or other transformation files in these public prefixes.
Questions
I would therefore be very grateful for clarification on the following points:
-
Were the 198,847 cells intentionally excluded from the released CCF coordinates as part of registration QC, tissue-artifact masking, or another quality-control step?
-
In particular, were the six complete sections
.01–.04and.68–.69
intentionally excluded from the final CCF coordinate release? -
What is the reason for the partial exclusion of 42,070 cells from the 11 otherwise registered sections?
-
The MERFISH-to-CCF registration documentation describes a global affine transformation followed by section-wise affine and deformable transformations.
Are the original transformation matrices / deformation fields used for specimenC57BL6J-638850available somewhere? -
If these transformation files are available, would it be methodologically appropriate to use them to transform the original coordinates of the 198,847 missing cells into CCFv3 coordinates?
-
Alternatively, if these cells were deliberately excluded because their registration was considered unreliable, would you recommend leaving them without CCF coordinates rather than attempting reconstruction?
Our goal is to retain the complete MERFISH cell-type and transcriptomic dataset while assigning CCF coordinates and Allen structure IDs only where the registration is considered reliable.
I have retained the complete list of the 198,847 missing cell_label identifiers and can provide section-specific summaries or the corresponding identifier table if this would help diagnose the issue.
Thank you very much for any clarification regarding the registration/QC procedure or availability of the original transformation files.
Best regards,
Oliver Schmitt

