Skip to content

Check the type of each CSV value instead of failing the read - #1724

Open
markhm wants to merge 1 commit into
datacontract:mainfrom
markhm:fix/csv-type-checks
Open

markhm wants to merge 1 commit into
datacontract:mainfrom
markhm:fix/csv-type-checks

Conversation

@markhm

@markhm markhm commented Oct 10, 2026

Copy link
Copy Markdown

Fixes #1721.

Problem

On a CSV file, DuckDB converts each column to the type its logicalType declares while it reads the file. The first value that does not convert, say two in an integer column, stops the run with Conversion Error: CSV Error on Line: 3, and every other check comes back without a result. A decimal in an integer column passes unnoticed, because CAST('2.50' AS BIGINT) rounds to 3. A date in another format than logicalTypeOptions.format declares either stops the run or passes.

What this changes

  • A CSV file is read with every column as text into the {model}__raw__ view. The typed table is then built from that text with TRY_CAST, so a bad value becomes NULL instead of stopping the read.
  • Each integer, number, boolean, date, timestamp and time column gets a field_type check, "Check that field quantity has type integer". It is an invalid_count check on the raw view, so it counts every bad value, and --include-failed-samples shows the values and their keys.
  • integer is strict: the text must be a whole number (^[+-]?[0-9]+$), so 2.50 and 2.00 fail.
  • date, timestamp and time with logicalTypeOptions.format must be in that format and must exist (31/02/2024 fails). The Java DateTimeFormatter pattern is mapped to a strptime format in the new csv_values.py. A pattern that cannot be mapped (for example XXX, which writes Z for UTC) gives a warning instead of a false failure. Without a format, ISO 8601 applies, as before.
  • The other checks keep running. A bad value is NULL for the other checks of its column, as for JSON files since 025d8ea, so a required column also reports it as missing.
  • Other formats and servers are unchanged.

Tests

  • tests/test_test_local_csv_types.py: 26 tests, covering a bad value of each type failing only its own column's check, every bad value counted, empty values, a bad value being missing for the other checks (also required), failed samples and diagnostics, ISO 8601 and declared date formats, the warning for a pattern that cannot be mapped, and the pattern mapping itself.
  • test_test_metadata_only.py and test_test_dimension_filter.py are updated for the new check, which is skipped with --metadata-only and belongs to the conformity dimension.
  • uv run pytest: all pass. uv run ruff check and ruff format: clean.

Docs

The data-types table in reference/local.md, the logicalTypeOptions.format bullet in schema.md, and the notes on CSV types in the s3, azure, gcs, index, testing/local.md and testing/xml.md pages. CHANGELOG entry under Fixed.

Known limit

--filter does not narrow the type check: it counts the whole file, because the raw view has no typed columns to filter on.

  • Tests pass (uv run pytest)
  • Code formatted (uv run ruff check --fix && uv run ruff format)
  • Docs updated (if relevant)
  • CHANGELOG.md entry added

🤖 Generated with Claude Code

A CSV file is read as text; each integer, number, boolean, date, timestamp
and time column is converted from it, and a type check counts the values
that do not convert. A bad value no longer stops the run with a conversion
error, and is missing for the other checks of its column. Integers must be
whole numbers, and dates, timestamps and times must be in their declared
format (ISO 8601 without one).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

CSV: a value that does not match its logicalType stops the run instead of failing a check

1 participant