PhilipGAQ commited on
Commit
0e1b7d3
·
verified ·
1 Parent(s): d13d545

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +126 -9
README.md CHANGED
@@ -3,14 +3,131 @@ license: cc-by-nc-sa-4.0
3
  language:
4
  - zh
5
  pipeline_tag: sentence-similarity
 
 
 
 
 
 
 
 
 
6
  ---
7
 
8
- @misc{jiang2026cmedtebcarebenchmarking,
9
- title={CMedTEB & CARE: Benchmarking and Enabling Efficient Chinese Medical Retrieval via Asymmetric Encoders},
10
- author={Angqing Jiang and Jianlyu Chen and Zhe Fang and Yongcan Wang and Xinpeng Li and Keyu Ding and Defu Lian},
11
- year={2026},
12
- eprint={2604.10937},
13
- archivePrefix={arXiv},
14
- primaryClass={cs.IR},
15
- url={https://arxiv.org/abs/2604.10937},
16
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  language:
4
  - zh
5
  pipeline_tag: sentence-similarity
6
+ library_name: transformers
7
+ tags:
8
+ - chinese
9
+ - medical
10
+ - information-retrieval
11
+ - dense-retrieval
12
+ - text-embeddings
13
+ - asymmetric-encoder
14
+ - cmedteb
15
  ---
16
 
17
+ # CARE-0.3B-4B
18
+
19
+ CARE-0.3B-4B is a complete asymmetric dense retrieval system for Chinese medical text retrieval. It consists of a lightweight 0.3B query encoder and a 4B document encoder.
20
+
21
+ The query encoder is intended for low-latency online query encoding, while the document encoder is intended for offline document encoding and indexing. The two encoders are trained as a pair and should be used together.
22
+
23
+ ## Model Components
24
+
25
+ | Component | Size | Recommended usage |
26
+ |---|---:|---|
27
+ | Query encoder | 0.3B | Online query encoding |
28
+ | Document encoder | 4B | Offline document encoding |
29
+
30
+ This repository/model release represents the paired `0.3B + 4B` CARE system. The 4B document encoder should not be paired with an unrelated query encoder when reproducing the reported results.
31
+
32
+ ## Intended Use
33
+
34
+ CARE-0.3B-4B is intended for:
35
+
36
+ - Chinese medical passage retrieval
37
+ - Medical knowledge-base search
38
+ - Retrieval-augmented generation over Chinese medical documents
39
+ - Offline document embedding and vector indexing
40
+ - Research on asymmetric dense retrieval
41
+
42
+ Typical deployment:
43
+
44
+ 1. Encode the document corpus offline with the 4B document encoder.
45
+ 2. Store document embeddings in a vector index.
46
+ 3. Encode incoming queries online with the 0.3B query encoder.
47
+ 4. Retrieve documents using dot-product or cosine similarity.
48
+
49
+ ## Out-of-Scope Use
50
+
51
+ This model is not a standalone chatbot, reranker, or medical diagnosis system. Retrieved passages must not be treated as medical advice or as a substitute for clinical judgment.
52
+
53
+ ## Inference
54
+
55
+ The inference wrapper is provided in [`inference/asymmetric.py`](https://github.com/PhilipGAQ/CARE/blob/main/inference/asymmetric.py).
56
+
57
+ ```python
58
+ from inference.asymmetric import CARE
59
+ import numpy as np
60
+
61
+ model = CARE(
62
+ model_name_or_path_query="path/to/CARE-0.3B-query-encoder",
63
+ model_name_or_path_doc="PhilipGAQ/CARE-0.3B-4B",
64
+ trust_remote_code=True,
65
+ use_fp16=False,
66
+ normalize_embeddings=True,
67
+ query_batch_size=2,
68
+ passage_batch_size=2,
69
+ )
70
+
71
+ queries = ["什么是高血压?"]
72
+ documents = [
73
+ "高血压是指动脉血压持续升高,通常指收缩压≥140mmHg和/或舒张压≥90mmHg。"
74
+ ]
75
+
76
+ query_embeddings = model.encode_queries(queries, task_name="retrieval")
77
+ document_embeddings = model.encode_corpus(documents, task_name="retrieval")
78
+
79
+ scores = np.dot(query_embeddings, document_embeddings.T)
80
+ print(scores)
81
+ ```
82
+
83
+ When `normalize_embeddings=True`, embeddings are L2-normalized and dot product is equivalent to cosine similarity. Similarity scores are ranking signals, not calibrated probabilities.
84
+
85
+ ## Training
86
+
87
+ CARE uses a two-stage asymmetric training strategy:
88
+
89
+ 1. Query-side alignment training with the document encoder fixed.
90
+ 2. Joint fine-tuning of the query and document encoders.
91
+
92
+ This progressively aligns representations produced by the structurally different query-side and document-side encoders.
93
+
94
+ ## Evaluation
95
+
96
+ The paired CARE system is evaluated on the Chinese Medical Text Embedding Benchmark (CMedTEB), which covers retrieval, reranking, and semantic textual similarity (STS).
97
+
98
+ Results should be reported for the complete configuration:
99
+
100
+ `CARE 0.3B query encoder + CARE 4B document encoder`
101
+
102
+ See the paper for benchmark splits, metrics, baselines, and full results.
103
+
104
+ ## Limitations and Responsible Use
105
+
106
+ - The model is primarily optimized for Chinese medical text retrieval.
107
+ - Performance may degrade on non-Chinese, non-medical, or highly specialized domains.
108
+ - Retrieval quality depends on chunking, preprocessing, document quality, and indexing strategy.
109
+ - Retrieved information may be incomplete, outdated, duplicated, or clinically inappropriate.
110
+ - The model does not verify clinical correctness and must not be used alone for diagnosis, treatment, medication, triage, or patient-specific risk decisions.
111
+ - Medical applications require qualified human review, source attribution, freshness checks, and appropriate privacy and safety controls.
112
+
113
+ ## Resources
114
+
115
+ - Code: https://github.com/PhilipGAQ/CARE
116
+ - Benchmark: https://huggingface.co/datasets/PhilipGAQ/CMedTEB
117
+ - Paper: https://arxiv.org/abs/2604.10937
118
+
119
+ ## Citation
120
+
121
+ ```bibtex
122
+ @inproceedings{jiang2026benchmarking,
123
+ title={Benchmarking and Enabling Efficient Chinese Medical Retrieval via Asymmetric Encoders},
124
+ author={Jiang, Angqing and Chen, Jianlyu and Wang, Yongcan and Li, Xinpeng and Ding, Keyu and Lian, Defu and others},
125
+ booktitle={Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
126
+ pages={20000--20020},
127
+ year={2026}
128
+ }
129
+ ```
130
+
131
+ ## License
132
+
133
+ CC-BY-NC-SA-4.0