Missing CCF coordinates for 198,847 cells in MERFISH-C57BL6J-638850 – availability of registration transforms?

Dear Allen Brain Cell Atlas / CCF registration team,

I am currently integrating the whole-brain MERFISH dataset MERFISH-C57BL6J-638850 with the Allen Mouse CCF and an existing mouse brain connectome/ontology database.

During quality control, I noticed a systematic difference between the complete MERFISH cell metadata and the publicly available CCF registration products.

For the original MERFISH dataset,

MERFISH-C57BL6J-638850/metadata/views/cell_metadata_with_cluster_annotation.csv

I obtain:

  • 3,938,808 unique cells

  • 59 brain sections

This agrees with the documented size of the MERFISH dataset.

However, the CCF registration products from

MERFISH-C57BL6J-638850-CCF/20231215

contain only 3,739,961 unique cells in each of the following files:

  • reconstructed_coordinates.csv

  • ccf_coordinates.csv

  • views/cell_metadata_with_parcellation_annotation.csv

Thus, 198,847 MERFISH cells (5.05%) are absent from all three CCF-derived products.

I compared the cell_label identifiers directly, so this difference does not result from duplicate cells or from a downstream join performed on our side.

1. Six complete sections are absent

The following six MERFISH sections are present in the original cell metadata but completely absent from the CCF-derived files:

  • C57BL6J-638850.01 — 23,389 cells

  • C57BL6J-638850.02 — 25,792 cells

  • C57BL6J-638850.03 — 30,546 cells

  • C57BL6J-638850.04 — 43,047 cells

  • C57BL6J-638850.68 — 20,040 cells

  • C57BL6J-638850.69 — 13,963 cells

Together these account for 156,777 cells.

2. Additional cells are missing from otherwise registered sections

Another 42,070 cells are absent from the CCF products although their respective sections are otherwise represented.

These missing cells occur in 11 sections.

Some examples are:

  • .27: 8,118 / 58,039 cells missing (13.99%)

  • .42: 8,051 / 75,551 (10.66%)

  • .33: 8,425 / 85,849 (9.81%)

  • .48: 5,915 / 77,165 (7.67%)

  • .24: 2,622 / 64,243 (4.08%)

  • .32: 36 / 93,525 (0.038%)

The missing cells appear to be spatially structured rather than randomly distributed.

For example, in section .27, the 1% x-coordinate quantile of the CCF-registered cells is approximately 0.96, whereas for the 8,118 missing cells it is approximately 6.80.

The missing populations are also not obviously characterized by poor transcriptomic annotation quality. Their mean average_correlation_score values are generally in the same range as the remaining dataset, and the missing cells comprise multiple classes and hundreds of clusters.

3. Public S3 resources

I also recursively inspected the public S3 resources for this dataset.

For

metadata/MERFISH-C57BL6J-638850-CCF/20231215/

I find only:

  • ccf_coordinates.csv

  • reconstructed_coordinates.csv

  • views/cell_metadata_with_parcellation_annotation.csv

For

image_volumes/MERFISH-C57BL6J-638850-CCF/20230630/

I find:

  • resampled_annotation.nii.gz

  • resampled_annotation_boundary.nii.gz

  • resampled_average_template.nii.gz

I could not find the global affine transform, section-wise affine transforms, deformation fields, displacement fields, or other transformation files in these public prefixes.

Questions

I would therefore be very grateful for clarification on the following points:

  1. Were the 198,847 cells intentionally excluded from the released CCF coordinates as part of registration QC, tissue-artifact masking, or another quality-control step?

  2. In particular, were the six complete sections
    .01–.04 and .68–.69
    intentionally excluded from the final CCF coordinate release?

  3. What is the reason for the partial exclusion of 42,070 cells from the 11 otherwise registered sections?

  4. The MERFISH-to-CCF registration documentation describes a global affine transformation followed by section-wise affine and deformable transformations.
    Are the original transformation matrices / deformation fields used for specimen C57BL6J-638850 available somewhere?

  5. If these transformation files are available, would it be methodologically appropriate to use them to transform the original coordinates of the 198,847 missing cells into CCFv3 coordinates?

  6. Alternatively, if these cells were deliberately excluded because their registration was considered unreliable, would you recommend leaving them without CCF coordinates rather than attempting reconstruction?

Our goal is to retain the complete MERFISH cell-type and transcriptomic dataset while assigning CCF coordinates and Allen structure IDs only where the registration is considered reliable.

I have retained the complete list of the 198,847 missing cell_label identifiers and can provide section-specific summaries or the corresponding identifier table if this would help diagnose the issue.

Thank you very much for any clarification regarding the registration/QC procedure or availability of the original transformation files.

Best regards,
Oliver Schmitt

Hi, Oliver

I can address question 1, 2. Hopefully, my colleague Michael Kunst can address your additional questions when he is back from leave. The cells from the six complete sections were excluded from the final CCF coordinate release is because our CCF reference do not contain the most anterior and posterior regions. These sections belong to the areas missing along the A-P axis in our CCF reference. This is a question we hope to address in future CCF release.

Hi,

(attachments)


Hi,

many thanks — this is very helpful and clarifies the issue.

If I understand correctly, the six complete sections were excluded from the final CCF-coordinate release not because of a problem with the underlying MERFISH data or section quality, but because they fall outside the anterior–posterior extent represented in the current CCF reference. In other words, the cells themselves are valid observations, but no corresponding CCF coordinates could be assigned within the present reference space.

This distinction is important for our analysis, because it means that we should treat these cells as outside the supported CCF domain rather than as biologically missing or failed data.

Thank you also for forwarding the remaining questions to Michael Kunst. I am very happy to wait until he is back from leave.

Best regards,
Oliver