现有数据标准的网站映射指南
本项目已经完成数据标准化。网站中的 JSON 仅用于公开检索、样本卡片与详情展示,是对现有标准的索引映射,不替代或重新定义原标准。真实原始数据、处理后体数据及其权威元数据仍按项目现有规范保存。
当前样本索引包含 6 张项目提供图像,全部为验证集。名称依据原文件名;图像像素保持原样,仅在网页展示时旋转、取景与适度调整亮度;完整原图保持原貌。体数据元数据保留空值,尚无体数据附件。
1. 接入顺序
- 确定已有标准的名称、版本、字段字典与维护来源;由项目负责人确认哪些字段可公开。
- 建立原字段 → 网站字段的对照,明确转换规则、单位和缺失值语义。必要时扩展网站映射层;不要为了迎合页面而改变已有科学含义。
- 复制
data/sample.template.json,为一条经批准公开的真实样本填写元数据。模板没有有效实验数值;其 is_demo: false 仅是填报真实条目的预设,不能证明它是真实样本,也不能直接发布。
- 对照
data/metadata.schema.json 检查类型、必填项与枚举,再验证实际文件和科学含义。Schema 校验通过不等于数据正确或获准公开。
- 将通过审核的条目放入
data/samples.json 的顶层数组,替换或清楚隔离 DEMO。仅在确为真实样本时使用 is_demo: false。
- 用 HTTP 服务预览并检查筛选、详情、单位、空值和资源状态,确认后再批量映射。
权威字段类型、层级和取值范围以当前 Schema 与模板为准;修改它们时也要同步 assets/js/app.js 和本指南。
当前 JSON 结构
samples.json 的根是样本数组,sample.template.json 的根是单个样本对象。metadata.schema.json 使用 JSON Schema Draft 2020-12 并校验数组,因此单条模板需要先包装为 [template] 再校验。不要把模板对象直接作为整个索引提交,也不要额外套上 samples 属性。
每条样本包含以下字段。列为“可空”的字段仍需保留字段名;准备阶段允许 null 不表示发布审核已经通过。
id |
非空字符串;只允许英文字母、数字、点、下划线、短横线,首位为字母或数字;发布时另查唯一性 |
title / description |
公开标题和说明;支持字符串或 { "zh": "中文", "en": "English" };标题不可为空 |
is_demo |
布尔值;合成界面演示为 true |
category |
phantom、simulation、in-vivo、ex-vivo 或 other |
split |
train、validation、test 或 unspecified |
tags |
无重复字符串的数组;无标签时用 [] |
preview |
项目相对路径或 HTTPS 图片地址,可空;示意图必须明确标注 |
dimensions |
三个正整数,按 axis_order 排列,可空 |
axis_order |
x、y、z 各出现一次的三元素数组,可空;不等同于解剖方向 |
spacing_mm |
三个正数,单位毫米,顺序与 axis_order 一致,可空 |
wavelength_nm |
单个正数,单位纳米,可空;多波长需明确扩展 Schema |
intensity_unit |
信号/强度单位的字符串或双语对象,可空 |
data_version |
该条目对应的数据版本字符串,可空 |
provenance |
必须保留的来源对象,见下方四项 |
files |
文件对象数组;未接入时为 [] |
provenance 必须包含四个可空字段:source_id(公开来源标识)、acquisition(阵列/采集配置来源)、reconstruction(重建方法与版本来源)、standardization_reference(已有标准及其版本的可追溯说明)。source_id 保持字符串;其余三个说明字段支持字符串或双语对象。当前没有单独的阵列几何对象;需要结构化扩展时同步 Schema 和页面。
每个 files 元素必须包含 name、format、url、size_bytes、sha256。其中 name 为非空字符串,format 为格式说明字符串;后面三项可为 null。非空 size_bytes 为非负整数,非空 sha256 为 64 位十六进制字符串。Schema 只检查这些基础约束,不自动保证 URL 可访问、单位正确、ID 唯一或资料获准公开。
2. 最低核对内容
身份与分组
- 稳定且唯一的
id、公开 title 和 description。
category、split、tags 应映射真实研究定义。当前六张图像均按项目方要求归入验证集;仿体分类依据文件名,其余不推断在体或离体。
- 如果真实类别或划分不在当前枚举中,扩展 Schema、界面选项和统计逻辑后再使用;不要把未知类别硬套到已有选项。
- 数据分组应遵从原研究方案,核查同一对象、实验或强关联样本跨集合的泄漏风险。
尺度、单位与坐标
- 数组尺寸只说明每一轴的元素数量,必须同时核实轴顺序和空间坐标约定。
spacing_mm 表达体素间距时,要确认与尺寸轴一一对应,并正确完成原单位到毫米的转换。
wavelength_nm 应来自实际记录;多波长资料不得因页面仅显示单个数值而丢失信息。
- 明确数值的物理单位、归一化状态及数据类型。任意单位与物理单位不能混用。
- 未知、不适用、未提供和零值含义不同,按 Schema 使用允许的缺失表达,不能用
0 代替未知值。
采集与重建来源
- 保留可公开的阵列类型/几何、采集配置、重建方法及版本来源。
- 原始信号、重建体数据和算法输出之间的关系可追溯。
- 来源字段用于说明可信出处,不应包含受试者身份、内部路径、密码或令牌。
- 不能公开的来源信息留在受控记录中;公开页面明确其可获得程度,避免伪造说明。
版本与资源
- 每条资源的版本与对应数据发布版本一致,并在内容变更时更新。
files 仅填写真实存在且经授权公开或受控分发的资源;没有可用文件时保留空数组。
- 下载 URL 必须实际验证;永久入口优于会过期的临时签名地址。
- 文件大小、格式和 SHA-256 从最终发布文件计算,不能由文件名猜测。
- 占位 URL、空的 DOI 和尚未选择的许可证保持空值,不创建看似可用的假链接。
3. 建议保留的映射记录
维护者可在自己的数据处理工程中保存下列对照;只有经审核的公开版本才应进入站点:
| 待填写 |
待核实 |
以 Schema 为准 |
如单位换算、轴重排 |
以原标准和 Schema 为准 |
与源记录/文件比对 |
“转换规则”应记录真实操作,尤其是坐标重排、重采样、波长选择和版本合并。若网站只做展示而不转换,明确记录为直接映射。
4. 大文件、链接与校验
本仓库只保存站点和轻量公开索引。大型扫描文件、压缩包和模型权重建议保存在项目批准的数据仓库或对象存储中,并在网站中链接。受控数据需要具备访问控制的服务;静态 Pages 页面及隐藏按钮不能保护数据。
对已准备的最终文件可在本地计算 SHA-256:
# Linux
sha256sum path/to/approved-file
# macOS
shasum -a 256 path/to/approved-file
把输出的真实哈希填入对应资源字段,并保留文件名、字节大小和版本记录。校验值用于完整性验证,不证明数据科学质量,也不替代发布许可。
上线前使用无登录的空白浏览器会话核实公共链接,检查最终下载文件、重定向、权限要求和校验值。不要只判断 HTTP 200,因为返回内容可能是登录页或错误页面。
5. 发布边界
- 公开索引自身也属于公开数据;不能夹带身份信息或敏感元数据。
- 对人体/动物资料及其他受限制数据,先核实审批、同意、许可和去标识化要求。
- 真正数据尚未接入时,保留页面 DEMO 标记及资源待提供状态。
- 样本数量、总容量和覆盖范围只能由已验证的真实发布清单得出。
- 任何文件、元数据或许可变更都应同步版本、引用和校验记录。
完成后继续逐项核对发布前检查清单。
6. 双语展示字段
title、description、intensity_unit 和 provenance 中的 acquisition、reconstruction、standardization_reference 支持原有字符串或 { "zh": "中文说明", "en": "English description" }。可空字段仍可为 null。双语对象至少提供 zh 或 en 中的一个字符串,建议补齐两种语言;标题中已提供的语言版本不可为空。运行时先显示当前语言,缺少有效文字时按中文、英文顺序回退。
id、category、split、tags、data_version、provenance.source_id、文件名、URL、数值和校验值保持原始格式。译文只能解释相同事实,不能修改数据、技术标识或演示状态。请核对两种语言均明确标注合成 DEMO,并将未提供或未经核验的研究信息保留为空或待提供。
Website mapping guide for the existing data standard
Data standardization for this project is already complete. The website JSON is only an index mapping for public search, sample cards, and detail views. It does not replace or redefine the original standard. Real raw data, processed volumes, and their authoritative metadata remain stored according to the project's existing specifications.
The sample index contains six project-supplied images, all in Validation. Names follow source filenames. Original image pixels are unchanged; the webpage adjusts their presentation rotation, framing and brightness, with an unfiltered full-original view. Volume metadata remains null and no volume attachments are available.
1. Integration order
- Establish the existing standard's name, version, field dictionary, and authoritative maintenance source. Have the project lead confirm which fields may be public.
- Map original fields → website fields, documenting conversion rules, units, and missing-value meanings. Extend the website mapping layer if necessary; do not change existing scientific meaning to fit the page.
- Copy
data/sample.template.json and enter metadata for one real sample approved for public release. The template contains no valid experimental values. Its preset is_demo: false is intended for preparing a real record; it is not proof of authenticity and does not make the template publishable.
- Check types, required fields, and enumerations against
data/metadata.schema.json, then validate the actual files and scientific meaning. Passing Schema validation does not establish correctness or permission to publish.
- Put approved records in the top-level array of
data/samples.json, replacing or clearly separating DEMO records. Use is_demo: false only for genuine real samples.
- Preview through an HTTP server and check filters, details, units, missing values, and resource states before mapping records in bulk.
The current Schema and template are authoritative for field types, nesting, and allowed values. When changing them, also update assets/js/app.js and this guide.
Current JSON structure
The root of samples.json is a sample array; the root of sample.template.json is one sample object. metadata.schema.json uses JSON Schema Draft 2020-12 and validates an array, so wrap a single template as [template] before validation. Do not submit the template object as the entire index or add an extra samples property around it.
Each sample has the following fields. Fields described as nullable must still retain their field names. Allowing null during preparation does not mean release review has passed.
id |
Nonempty string; letters, digits, periods, underscores, and hyphens only, starting with a letter or digit. Check uniqueness separately before release. |
title / description |
Public title and description; accept strings or bilingual { "zh": "中文", "en": "English" } objects. Titles must not be empty. |
is_demo |
Boolean; true for synthetic interface demonstrations. |
category |
phantom, simulation, in-vivo, ex-vivo, or other. |
split |
train, validation, test, or unspecified. |
tags |
Array of unique strings; use [] when there are no tags. |
preview |
Project-relative path or HTTPS image URL; nullable. Clearly label illustrations. |
dimensions |
Three positive integers in axis_order; nullable. |
axis_order |
Three-element array containing x, y, and z once each; nullable. This does not specify anatomical orientation. |
spacing_mm |
Three positive numbers in millimeters, in the same order as axis_order; nullable. |
wavelength_nm |
One positive number in nanometers; nullable. Multiple wavelengths require an explicit Schema extension. |
intensity_unit |
Signal/intensity unit string or bilingual object; nullable. |
data_version |
Data version string for the record; nullable. |
provenance |
Required provenance object containing the four fields below. |
files |
Array of file objects; use [] until resources are integrated. |
provenance must contain four nullable fields: source_id (public source identifier), acquisition (source for the array/acquisition configuration), reconstruction (source for the reconstruction method and version), and standardization_reference (traceable reference to the existing standard and version). source_id remains a string; the other three description fields accept strings or bilingual objects. There is currently no separate array-geometry object. Update the Schema and page together if a structured extension is needed.
Each files element must contain name, format, url, size_bytes, and sha256. name is a nonempty string and format is a format-description string; the last three may be null. A non-null size_bytes is a nonnegative integer, and a non-null sha256 is a 64-character hexadecimal string. The Schema checks only these basic constraints. It does not ensure that URLs work, units are correct, IDs are unique, or publication is authorized.
2. Minimum checks
Identity and grouping
- A stable, unique
id, public title, and description.
category, split, and tags must map to actual research definitions. All six images are assigned to Validation by the project owner. Phantom categories follow filenames; in-vivo or ex-vivo status is not inferred.
- If real categories or splits fall outside current enumerations, extend the Schema, interface options, and statistics logic first. Do not force unknown categories into existing choices.
- Follow the original study design for data splitting. Check for leakage when the same subject, experiment, or strongly related samples occur across sets.
Scale, units, and coordinates
- Array dimensions state only the element count along each axis. Verify axis order and spatial coordinate conventions as well.
- When
spacing_mm represents voxel spacing, ensure its entries correspond to the dimension axes and correctly convert original units to millimeters.
- Obtain
wavelength_nm from actual records. Do not discard multiwavelength information because a page displays only one number.
- Specify physical units, normalization status, and data type. Arbitrary units and physical units are not interchangeable.
- Unknown, not applicable, not provided, and zero have different meanings. Use missing-value representations permitted by the Schema; do not use
0 for unknown values.
Acquisition and reconstruction provenance
- Retain publicly shareable sources for array type/geometry, acquisition configuration, reconstruction method, and version.
- Keep relationships between raw signals, reconstructed volumes, and algorithm outputs traceable.
- Provenance fields should identify trusted sources. They must not include participant identities, internal paths, passwords, or tokens.
- Keep nonpublic provenance in controlled records. State its availability clearly on public pages rather than fabricating descriptions.
Versions and resources
- Each resource version must match its associated data release and be updated when content changes.
- Include in
files only existing resources authorized for public or controlled distribution. Keep an empty array when no files are available.
- Test download URLs in practice. Prefer permanent entry points to expiring signed URLs.
- Compute file size, format information, and SHA-256 from the final release files; do not infer them from filenames.
- Leave placeholder URLs, absent DOIs, and undecided licenses empty. Do not create apparently usable fake links.
3. Recommended mapping record
Maintainers can keep the following record in their own data-processing project. Only an approved public version should enter the website:
| To be supplied |
To be verified |
Follow the Schema |
For example, unit conversion or axis reordering |
Follow the original standard and Schema |
Compare with source records/files |
Record actual operations under “Conversion rule,” especially coordinate reordering, resampling, wavelength selection, and merging versions. If the website only displays data without conversion, explicitly record a direct mapping.
4. Large files, links, and checksums
This repository holds only the website and lightweight public index. Store large scan files, archives, and model weights in a project-approved data repository or object store, then link to them. Controlled data needs a service with access control. Static Pages hosting and hidden buttons cannot protect data.
Compute SHA-256 locally for prepared final files:
# Linux
sha256sum path/to/approved-file
# macOS
shasum -a 256 path/to/approved-file
Enter the actual hash in the corresponding resource field, retaining filename, byte size, and version records. Checksums verify integrity; they do not establish scientific quality or replace publication permission.
Before launch, test public links in a fresh, signed-out browser session. Check the final downloaded files, redirects, access requirements, and checksums. An HTTP 200 response alone is insufficient because the content may be a sign-in or error page.
5. Release boundaries
- The public index is itself public data. It must not contain identifying information or sensitive metadata.
- For human, animal, or other restricted data, first verify approval, consent, licensing, and de-identification requirements.
- Retain DEMO labels and pending-resource states while real data has not been integrated.
- Derive sample counts, total size, and coverage only from a verified real release manifest.
- Update versions, citations, and checksum records whenever files, metadata, or licenses change.
Then work through the pre-release checklist.
6. Bilingual display fields
title, description, intensity_unit, and provenance fields acquisition, reconstruction, and standardization_reference accept a legacy string or { "zh": "中文说明", "en": "English description" }. Nullable fields may still be null. A bilingual object must provide at least one zh or en string; supplying both is recommended. Every supplied title translation must be nonempty. The runtime displays the active language first, falling back to Chinese and then English when valid text is missing.
Keep id, category, split, tags, data_version, provenance.source_id, filenames, URLs, numeric values, and checksums in their original formats. Translations must describe the same facts without changing data, technical identifiers, or demo status. Check that both languages clearly identify synthetic DEMO content and keep missing or unverified research information empty or pending.