描述
包含所有已缓存文件 schema 的相关信息。
列
storage(String) — 存储名称:File、URL、S3 或 HDFS。source(String) — 文件来源。format(String) — 格式名称。additional_format_info(String) — 用于识别 schema 所需的附加信息。例如,格式特定的设置。registration_time(DateTime) — schema 添加到缓存中的时间戳。schema(Nullable(String)) — 已缓存的 schema。number_of_rows(Nullable(UInt64)) — 文件在给定格式下的行数。它用于缓存数据文件的简单count()结果,以及在推断 schema 期间缓存元数据中的行数。schema_inference_mode(Nullable(String)) — schema 推断模式。
示例
假设我们有一个 data.jsonl 文件,内容如下:
{"id" : 1, "age" : 25, "name" : "Josh", "hobbies" : ["football", "cooking", "music"]}
{"id" : 2, "age" : 19, "name" : "Alan", "hobbies" : ["tennis", "art"]}
{"id" : 3, "age" : 32, "name" : "Lana", "hobbies" : ["fitness", "reading", "shopping"]}
{"id" : 4, "age" : 47, "name" : "Brayan", "hobbies" : ["movies", "skydiving"]}打开 clickhouse-client 并运行 DESCRIBE 查询:
DESCRIBE file('data.jsonl') SETTINGS input_format_try_infer_integers=0;┌─name────┬─type────────────────────┬─default_type─┬─default_expression─┬─comment─┬─codec_expression─┬─ttl_expression─┐
│ id │ Nullable(Float64) │ │ │ │ │ │
│ age │ Nullable(Float64) │ │ │ │ │ │
│ name │ Nullable(String) │ │ │ │ │ │
│ hobbies │ Array(Nullable(String)) │ │ │ │ │ │
└─────────┴─────────────────────────┴──────────────┴────────────────────┴─────────┴──────────────────┴────────────────┘来看看 system.schema_inference_cache 表中的内容:
SELECT *
FROM system.schema_inference_cache
FORMAT VerticalRow 1:
──────
storage: File
source: /home/droscigno/user_files/data.jsonl
format: JSONEachRow
additional_format_info: schema_inference_hints=, max_rows_to_read_for_schema_inference=25000, schema_inference_make_columns_nullable=true, try_infer_integers=false, try_infer_dates=true, try_infer_datetimes=true, try_infer_numbers_from_strings=true, read_bools_as_numbers=true, try_infer_objects=false
registration_time: 2022-12-29 17:49:52
schema: id Nullable(Float64), age Nullable(Float64), name Nullable(String), hobbies Array(Nullable(String))