All cows developed chronic mastitis infections, confirmed by bacterial plating and somatic cell counts (SCC), following S. aureus challenge with strain Newbould 305. Post-infection, animals averaged 2.82 × 106 ± 1.08 × 106 somatic cells/mL of the infected quarter. Only infected animals generated sufficient SCC for cell characterization as the NADC healthy herd SCC average was 2.1 × 104 somatic cells/mL. Leukocyte cells from blood and milk samples were collected and analyzed via scRNA-seq, resulting in cell clusters of all major immune cell types.
In total, 35,338 cells were recovered via scRNA-seq of milk (20,233 total cells) and blood (15,105 total cells) samples collected from each of three Holstein cattle with chronic mastitis (6 samples total; Fig. 1A). Sixty two (62) cell clusters with distinct transcriptional profiles were resolved (see methods; Fig. 1B) with identification of > 4.3 million differentially expressed genes recovered from all pairwise cluster comparisons (Supplementary Fig. 1). Despite their transcriptional diversity amongst the 60 + cell clusters, clusters were assigned to one of six major immune cell types (Fig. 1C) based on conserved, cell type-specific gene expression profiles, including (1) granulocytes (22,020 cells); (2) monocyte/macrophage/conventional dendritic cells (cDCs; (3,187 cells); (3) B cells and antibody-secreting cells (ASCs; (4,076 cells); (4) T cells and innate lymphoid cells (ILCs; (5,473 cells); (5) plasmacytoid dendritic cells (pDCs; (260 cells); and (6) non-immune cells (322 cells). The conserved canonical genes used to assign cell type identities were investigated through general assessment of expression (Fig. 1D,E) and query of genes differentially expressed in each cluster (Supplementary Table 1) to assign cell type identities as described below.
Single-cell RNA sequencing of milk and blood cells from cattle with chronic mastitis reveals six major cell types. (A–C) t-SNE plots showing sample origin (A), cluster assignment (B), and cell type annotation (C) of cells recovered via scRNA-seq. (D) t-SNE plots showing expression of a subset of canonical genes used to annotate cell types shown in (C). (E) Dot plot of canonical genes (x-axis) used to annotate clusters into cell types (y-axis). Dot size within the plot indicates the percentage of cells in a cluster expressing a gene. Dot fill color indicates the relative expression level of a gene in a cluster. The color bar on the left of the y-axis indicates cell type assignment of clusters. (F) t-SNE plots showing number of UMIs (top) or number of genes (bottom) expressed in each cell. In all t-SNE plots (A–D,F), each dot in the plot corresponds to an individual cell. Color of each dot in a plot corresponds to assignment of a cell to sample origin (A), cluster assignment (B), cell type annotation (C), level of gene expression (D), or number of UMIs or genes recovered from a cell (F). Proximity of cells within a plot is not necessarily a correlate of relatedness. *Gene name is not listed in the genome annotation file but was identified through manual query. See methods section “Gene name replacement” for further information. ASC antibody-secreting cell, cDC conventional dendritic cell, ILC innate lymphoid cell, pDC plasmacytoid dendritic cell, scRNA-seq single-cell RNA sequencing, t-SNE t-distributed stochastic neighbor embedding, UMI unique molecular identifier.
Granulocytes and monocytes/macrophages/cDCs shared expression of many myeloid lineage genes, including CSF3R, CXCL8, SRGN, TYROBP, SIRPA, CD68, TLR4, ADGRE1, and HCK9,15,20,21,22. However, granulocytes lacked expression of other genes used as key markers of monocytes/macrophages/cDCs, including CD14, TREM2, AIF1, CSF1R, CST3, CD86, CD83, FCGR3A, FLT3, IRF8, BOLA-DQA2and BOLA-DRB39,15,20,21. BOLA-DQA2 and BOLA-DRB3 were not listed in the genome annotation file but were identified through manual query (see methods section “Gene name replacement” for further information). Moreover, granulocyte clusters had lower transcriptional complexity in terms of both UMIs captured per cell and genes captured per cell, similar to previous reports (Fig. 1F)23. B cells and ASCs expressed genes encoding for components associated with B cell receptor (BCR) signaling (CD79A, CD79B, CD19, MS4A1, BLNK), antibody secretion (JCHAIN), and B lineage-defining transcription factors (PAX5, IRF4)9,15,20,21,24. T cells and ILCs were annotated based on expression of genes encoding for proteins associated with T cell receptor (TCR) complex signaling (CD3E, CD3D, CD3G, CD247, ZAP70), T cell subset markers (TRDC*, CD4, CD8A, CD8B), and markers characteristically expressed by natural killer (NK) cells, other ILC subsets, and/or cytotoxic T cells (NCR1, KLRB1, NKG2A, KLRD1, NKG7, GNLY, CTSW, PRF1)9,15,20,21. pDCs also expressed CD4 and further expressed additional pDC marker genes, including GZMB and IL3RA25. Non-immune cells were identified due to high expression of LALBA and CSN2, markers characteristic of bovine mammary epithelial cells15. Non immune cells also had very few cells expressing PTPRC (encoding CD45), a pan-leukocyte marker typically expressed by immune cells recovered via scRNA-seq20. Genes characteristic of erythrocytes were also queried but found to be absent (AHSP, HBB, HBM)21, indicating little to no red blood cell contamination in the dataset.
Milk somatic cells are dominated by granulocytes, most-so during infectious states where neutrophils undergo diapedesis to enter the mammary gland from the blood4 leading to vastly different cell profiles between blood and milk samples. Therefore, we initially investigated the differences in cell compositions recovered via scRNA-seq from milk and blood samples of cattle with chronic mastitis (Fig. 2A). Cells with high degrees of transcriptional similarities were allocated into cell neighborhoods, and within each cell neighborhood, the sample origin of each cell was determined and used to test for statistical differences in abundance for milk versus blood samples. Differential abundance testing resulted in identification of cell neighborhoods that were significantly enriched in milk or blood samples as shown in Fig. 2B. Cell neighborhoods were further assigned back to cell annotations at the cluster (Fig. 2C) or cell type (Fig. 2D) level. Within a cell cluster, significant differential abundance was exhibited in at least one cell neighborhood, with the exception of c18, c28, c34, c43, c56, and c58 that did not have any neighborhoods with significant differential abundance detected (Fig. 2C). Cluster-level findings indicated neighborhoods exhibited significantly greater abundance in blood for B/ASC, T/ILC, and pDC clusters, significantly greater abundance in milk for non-immune clusters, and a mixture of significantly greater abundance in milk versus blood for granulocytes and monocyte/macrophage/cDCs (Fig. 2C). Figure 2D summarizes the previously made observations at the cell type level, again showing significant differential abundance biased towards blood in B/ASCs, T/ILCs, and pDCs; biased towards milk in non-immune cells; and biased towards both milk and blood for granulocytes and monocyte/macrophage/cDCs. Overall, these findings could also be generally observed in the cell compositions obtained from each cluster or cell type, as shown in Fig. 2E and F, respectively.
Cell types exhibit different abundances in milk versus blood samples recovered from cattle with chronic mastitis. (A) t-SNE plot showing sample type cells were recovered from for scRNA-seq. Each dot in the plot corresponds to an individual cell. Color of each dot in the plot corresponds to assignment of a cell to milk (blue) or blood (red) samples. Proximity of cells within a plot is not necessarily a correlate of relatedness. (B) Results of differential abundance testing for cell neighborhoods overlaid onto t-SNE coordinates shown in (A). Each dot in the plot represents one cell neighborhood. Size of the dot corresponds to the number of cells in a neighborhood. Width of lines connecting neighborhoods corresponds to the number of cells shared between neighborhoods. The color of a neighborhood corresponds to significantly greater abundance in blood (red), milk (blue), or no significant difference in abundance (grey). Shade of red (blood) or blue (milk) fill in cell neighborhoods corresponds to the logFC in abundance. A neighborhood was considered differentially abundant if the spatially-corrected p-value was < 0.05. (C,E) Beeswarm plots of differential abundance results shown in B separated for individual cell clusters (C) or cell types (E) (y-axes). A neighborhood was assigned to a cluster (C) or cell type (E) if > 90% of cells in a neighborhood belonged to a single cluster. Differential abundance testing results for neighborhoods with ≤90% of cells recovered from a single cluster (C) or cell type (E) are not shown. Each dot in the plot represents one cell neighborhood. The color of a dot corresponds to significantly greater abundance in blood (red), milk (blue), or no significant difference in abundance (grey). Position on the x-axes and shade of red (blood) or blue (milk) fill in dots corresponds to the logFC difference in abundance. In C, the color bar to the left of the plot corresponds to the cell type each cluster was assigned to. To the right of each plot, the number of neighborhoods exhibiting significantly higher abundance in blood, no significant difference in abundance, and significantly higher abundance in milk are sequentially listed for each cluster (C) or cell type (F). A neighborhood was considered differentially abundant if the spatially-corrected p-value was < 0.05. (D,F) Stacked bar plots showing the proportion of cells (x-axes) comprising each cluster (D) or cell type (F) (y-axes). Fill color of bars corresponds to sample cells were derived from. In (D), the color bar to the left of the plot corresponds to the cell type each cluster was assigned to. Bar size is proportional to 100% of cells comprising a cluster (D) or cell type (F) and is not proportional to the total absolute number of cells in each group. ASC antibody-secreting cell, cDC conventional dendritic cell, ILC innate lymphoid cell, logFC log fold change, Nhood neighborhood, NS not significant, pDC plasmacytoid dendritic cell, scRNA-seq single-cell RNA sequencing, t-SNE t-distributed stochastic neighbor embedding.
Cell compositions for general cell types obtained via scRNA-seq between milk and blood samples were further validated by flow cytometry experiments conducted on milk and blood samples. Monocyte/macrophage, combined lymphocyte, B, T/ILC populations, and granulocyte cell type proportions determined through scRNA-seq analysis (Fig. 3A) were highly similar to those identified by flow cytometry (Fig. 3B). As reported previously, milk cells from infected quarters were dominated by granulocytes26, but populations of macrophages (cells expressing CD14), lymphocytes, and T cells (lymphocytes expressing CD3) were also present in similar proportions between scRNA-seq and flow cytometry data of milk.
Similar cell type proportions recovered from milk and blood via single-cell RNA sequencing and flow cytometry. (A) Stacked bar plot showing the proportion of cells (y-axis) comprising each sample used for scRNA-seq (x-axis). Fill color of bars corresponds to cell type. Bar size is proportional to 100% of cells comprising a sample and is not proportional to the total absolute number of cells in each sample. (B) Comparative cell type percentages as indicated by flow cytometry analysis. Percentages are relative to CD45 + live single cells with the exception of CD3 + cells whose percentages are within CD45 + live single cell lymphocytes (black arrow). ASC antibody-secreting cell, cDC conventional dendritic cell, ILC innate lymphoid cell, logFC log fold change, pDC plasmacytoid dendritic cell, scRNA-seq single-cell RNA sequencing.
Of exceptional interest in the context of mastitis infection were differentially expressed genes amongst granulocyte clusters (total of 8981 DEGs detailed in Supplemental Table 2). Further investigation was conducted to uncover potential gene signatures that may be unique to milk samples and thus indicative of a granulocyte-specific, localized immune response to mastitis. Hierarchical clustering of all granulocyte clusters revealed early transcriptional divergence of c30 and c46 into a separate node at a high level of classification (Fig. 4A). Assessment of the cellular compositions of high-level nodes revealed Node 1 (containing c30 and c46) was almost exclusively comprised of milk-derived cells (96.2%), while other nodes (Node 2, Node 3) contained a mixture of both milk and blood cells (Fig. 4B). A Node 1-enriched gene signature was created by identifying all genes with significantly greater expression in both c30 and c46 relative to other granulocyte clusters, yielding a list of 210 genes (Fig. 4C, blue box & Figure S2). Gene set enrichment analysis of the 210-gene signature revealed granulocyte clusters comprised of higher percentages of milk cells generally had higher average enrichment scores for the gene signature (Fig. 4D). In fact, the top four clusters with the highest percentage of milk cells (> 88% per cluster) also had the four highest average gene set enrichment scores (Fig. 4D, green box). Closer dissection of cells derived from milk versus blood samples within each granulocyte cluster demonstrated higher enrichment of the Node 1 gene signature for milk-derived cells compared to blood-derived cells across every granulocyte cluster, indicating the 210-gene signature could be ubiquitously applied to granulocyte clusters as a milk-enriched transcriptional program (Fig. 4E). Thus, the 210-gene milk-enriched granulocyte signature serves as an indicator of granulocyte-specific transcriptional programming of the localized immune response to mastitis. To determine functional implications of the milk-enriched granulocyte transcriptional program, the 210-gene signature was used as input for identification of enriched Gene Ontology (GO). The gene signature was significantly enriched for 65 biological process GO terms that clustered into 9 functional groups (Fig. 4F), the details of which are available in Supplemental Tables 3, and include terms such as ‘innate immune response’, ‘granulocyte chemotaxis’, ‘response to bacterium’, ‘toll-like receptor signaling pathway’, and ‘innate immune response-activating signal transduction’.
A milk-enriched granulocyte gene signature indicates transcriptional dyanmics of the granulocyte-specific, localized immune response to mastitis. (A) Phylogenetic tree indicating the transcriptional relatedness of granulocyte clusters recovered via scRNA-seq. The red dotted line indicates high-level segregation of hierarchical clustering into nodes referenced in the results. (B) Pie charts showing the proportion of cells (pie slices) comprising the cells within each node identified in (A). Fill color of pie slices corresponds to sample identifiers. Pie size is proportional to 100% of cells comprising a node and is not proportional to the total absolute number of cells in each node. (C) Dot plot of signature genes (x-axis) and their expression patterns across granulocyte clusters (y-axis). Signature genes were identified for node 1 identified in (A). The full node 1 gene set (210 genes) is not shown and can instead be found in Supplementary Fig. 2. Only genes with log2FC > 1 for both clusters (c30, c46) are shown in this figure. Dot size within the plot indicates the percentage of cells in a cluster expressing a gene. Dot fill color indicates the relative expression level of a gene in a cluster. The phylogenetic tree on the left shows relatedness of granulocyte clusters, identical to as shown in (A). (D) Left: Stacked bar plot showing the percentage of cells (x-axis) comprising each granulocyte cluster used for scRNA-seq (y-axis). Fill color of bars corresponds to cells derived from blood (red) or milk (blue) samples. Bar size is proportional to 100% of cells comprising a cluster and is not proportional to the total absolute number of cells in each cluster. Clusters on the y-axis are listed (top to bottom) in ascending order of the percentage of milk cells comprising each cluster. Right: heatmap showing the average gene set enrichment score for the 210-gene node 1 gene signature. Fill color corresponds to the average gene set enrichment score for each granulocyte cluster. (E) Violin plots showing gene set enrichment scores (y-axis) for the 210-gene node 1 gene signature. Granulocyte clusters (x-axis) are separated into cells derived from blood samples (red) or milk samples (blue). c60 does not have a violin plotted for blood since no blood-derived cells were identified in the cluster. Individual points within each violin represent single cells. (F) Cytoscape Gene Ontology (GO) term enrichment and clustering analysis of the 210-gene node 1 gene signature. scRNA-seq single-cell RNA sequencing.



