Files
PgTidy/docs/todo.md
T
Hein 41fdaf415c
CI / Test (push) Successful in 45s
CI / Build (push) Successful in 21s
feat(format): implement subquery and CASE expression formatting
* Add support for formatting subqueries with configurable placement and spacing.
* Implement CASE expression formatting with options for wrapping and collapsing.
* Introduce tests for subquery and CASE expression scenarios to ensure correctness.
2026-09-21 17:10:49 +02:00

289 lines
19 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# PgTidy — TODO / Progress
Tracks what is done and what remains. See `plan.md` for the full design and rationale.
Legend: ✅ done · 🚧 in progress · ⬜ not started
---
## V1 — Formatter + CLI (current milestone)
### ✅ Scaffold module + repo hygiene
- `go.mod` (`module git.warky.dev/wdevs/pgtidy`, go 1.26).
- Replaced WkMailSync boilerplate: `AGENTS.md` (PgTidy architecture + invariants),
`CLAUDE.md`, `Makefile` (`build`/`test`/`vet`/`fmt`/`lint`/`clean`).
- Directory layout created (`cmd/`, `pkg/...`, `testdata/`).
- Copied 4 real-world procedures into `testdata/corpus/*.pgsql` (safety harness).
### ✅ Lossless lexer — `pkg/lexer`
- `token.go`: `Kind` enum (trivia, words, literals, operators, punctuation) + `Token`.
- `lexer.go`: full PG token coverage — dollar-quoted strings (`$tag$`, nested-tag aware),
`--` / nestable `/* */` comments, standard/escape/bit/hex/unicode strings,
numbers (decimal, exponent, leading dot, `0x/0o/0b`, `_` separators), positional params
(`$1`), operator runs with PostgreSQL's trailing `+/-` rule, `::` / `:=` / `:`, punctuation.
Tracks line/col per token.
- `lexer_test.go`: unit tests (kinds, operator trailing rule, dollar quotes, line/col) +
**corpus round-trip** asserting `emit(Lex(src)) == src` byte-for-byte.
- **Status:** all tests pass; round-trips all 4 corpus files. Invariant #1 satisfied.
### ✅ CST node model + parser — `pkg/cst`, `pkg/parser`
- `pkg/cst`: lossless node model — `Tok` (significant token + leading trivia), `Attach`,
`File.Source()` (byte-exact reconstruction), `Raw` (verbatim fallback), `CreateFunction`
(Head/Name/Params/Options/As/Body/Tail/Semi), `Param` (with separator comma).
- `pkg/parser`: statement splitting at top-level `;`; structures CREATE FUNCTION/PROCEDURE
(head, qualified name, comma-split params, option clauses split by keyword, AS + body);
everything else → `Raw`. Graceful degradation is a property of the data model.
- `parser_test.go`: small round-trips, CreateFunction shape assertions, Raw fallback, and
**corpus round-trip** — reconstructs all 4 files byte-for-byte; structures all 4 functions.
- **Status:** all tests pass.
- _Still TODO (later): DML/other-DDL structuring (currently Raw) for full formatting._
### 🚧 PL/pgSQL body parser ← NEXT (the remaining V1 piece)
#### ✅ DECLARE section — `pkg/format/body.go`
- `formatBody` splits the dollar-quote tag, calls `formatBodyInner`.
- `formatBodyInner` locates `DECLARE` and `BEGIN` at depth 0, formats the declare block,
then emits `BEGIN` onwards verbatim.
- `formatDeclareVars`: each variable declaration collapsed to one line
(` name type [= expr];`), `--Block--` comment markers preserved on their own lines,
mid-declaration block comments trigger verbatim fallback.
- `needSpace` fixed for `LBracket` — no space before `[` after ident/closing bracket
(fixes `citext[]`, array subscripts).
- `semanticallyEqual` in tests updated to recurse into dollar-quoted body tokens so
whitespace normalization inside the body does not falsely fail the semantic check.
- Added `testdata/corpus/test_a_broken.pgsql` — a CRLF corpus file with intentionally
broken layout (split-line variables + split-line body statements) used as a formatting
target; `testdata/corpus/test_a.pgsql` is the golden output.
- **Status:** all tests pass; DECLARE section formats correctly.
#### ✅ Body statement formatter — `pkg/format/body.go`
- `formatBodyStatements`: line-by-line formatter for the `BEGIN … END` block.
- Block-depth tracker: `BEGIN`/`END`, `IF`/`THEN`/`ELSIF`/`ELSE`/`END IF`, `LOOP`/`END LOOP`, `EXCEPTION`.
- Split-line join: col-0 lines at paren-depth 0 with non-clause first token are joined to preceding line. SQL clause keywords (`SELECT`, `FROM`, `WHERE`, `INTO`, `WITH`, `HAVING`, `GROUP`, `ORDER`, `RETURNING`, `SET`, joins) stay on own lines.
- Base-indent normalisation: first logical line of each statement gets `blockDepth × st.Indent`; subsequent lines preserve their original indentation (relative indentation maintained for multi-line expressions).
- Blank-line count preservation: blank lines between statements kept as-is.
- Verbatim-indent mode after `EXCEPTION`: original leading whitespace preserved to avoid style conflicts between functions that put `WHEN` at col-0 vs indented.
- `sqlClauseKw` map; `firstBodyKeyword`, `leadingWhitespace` helpers.
- `TestFormatBodyBroken`: golden-file test — `format(test_a_broken.pgsql)` must equal `test_a.pgsql`.
- Updated `test_a.pgsql` to match actual formatter output.
- **Status:** all tests pass; idempotence verified.
### ✅ Printer + style config — `pkg/format`, `pkg/config`
- `pkg/config`: `Style` struct + `Default()` = house style. `Load(startDir)` walks up the
directory tree to find `.pgtidy.yaml` and merges its fields over the defaults.
Supported keys: `indent`, `newline`, `keyword_case`, `ident_case`, `type_case`, `commas`.
Dependency: `gopkg.in/yaml.v3`.
- `pkg/format`: formats CREATE FUNCTION/PROCEDURE **headers** to house style (params
one-per-line leading-comma, option clauses each on own line, AS/`$$` own lines); DECLARE
section formatted (see body parser entry); `Raw` statements emitted verbatim.
Spacing engine (`needSpace`, tight ops `:: : -> ->>`, array `[]`) + casing
(`keywords`/`typeNames` sets). Comment-safety: verbatim fallback if a header carries
comments it cannot relocate.
- `pkg/format/body.go`: DECLARE section formatter (see body parser entry).
- Tests: golden header, idempotence, corpus idempotence + **semantic equivalence**
(updated to recurse into dollar-quoted body tokens).
- _Note: not a full Wadler Doc-IR yet — fixed-layout printer. Doc-IR for width-based
expression wrapping can come when DML structuring lands._
### ✅ CLI `fmt` + safety harness — `cmd/pgtidy`
- `pgtidy fmt` (gofmt model): default stdin→stdout; `-w`/`--write`, `-l`/`--list`,
`--check` (CI exit codes); `-d`/`--diff` (unified diff output); `version`/`help`.
- Config discovery: walks up from cwd to find `.pgtidy.yaml`; applied before formatting.
- `diff.go`: in-house unified diff (LCS-based, zero additional deps).
- `fmt_test.go`: stdin, --check (un/formatted), -w idempotence, unknown-command.
- Safety invariants #2 (semantic equivalence), #3 (idempotence), #4 (graceful degradation)
are tested in `pkg/format` over the corpus.
- **Runtime safety gate** (`format.VerifySafe`, `pkg/format/safety.go`): every frontend runs
it before emitting. Bundles `SemanticallyEqual` (code token stream) + `CommentsPreserved`
(no comment dropped/merged/split/reworded, recursing into bodies) + `StructurallyBalanced`
(`()[]` / BEGIN·CASE·IF·LOOP…END profile, ignoring comment & string contents) + an
idempotence re-format. On failure the CLI prints the reason and keeps the original.
Caught two real bugs: multi-line `/* */` bodies were reindented as code, and col-0 `--`
lines were glued onto the previous line (merging consecutive comments) — both fixed in
`formatBodyStatements`.
---
## ✅ V2 — Linter
- `pkg/pgast`: `go-pgquery` (WASM, no cgo) wrapper → real PG AST. `FirstTokenOffset`
skips leading whitespace/comments for accurate line numbers.
- `pkg/diagnostics`: `Diagnostic{RuleID, Severity, Message, File, Line, Col}`.
- `pkg/lint`: `Engine`, `Rule` interface, `New()` with all built-ins:
- MIG001 CREATE INDEX without CONCURRENT
- MIG002 ALTER TABLE ADD COLUMN NOT NULL without DEFAULT
- MIG003 ALTER TABLE ADD CONSTRAINT FK/CHECK without NOT VALID
- COR001 SELECT * | COR002 UPDATE without WHERE | COR003 DELETE without WHERE
- NAM001/2/3 table/column/function names not snake_case (quoted identifiers only)
- `pgtidy lint [--only=ID,...] [files...]`; exits 1 on findings, 2 on error.
- Fixture SQL in `testdata/lint/`; 6 tests covering violations + clean fixtures.
- `--fix` rewrites files in place applying autofixes; for stdin, prints fixed SQL to stdout.
- Autofixable: **MIG001** (insert `CONCURRENTLY` after `INDEX`) and **MIG003** (insert `NOT VALID` before `;`). MIG002, COR*, NAM* are intentionally not autofixable.
- `pkg/diagnostics.TextFix{Offset, End, New, Title}` — byte-range replacement attached to `Diagnostic.Fix`.
- `pkg/lint.ApplyFixes` — applies all fixes in reverse-offset order; overlapping fixes skipped.
- Fix helpers (`mig001Fix`, `mig003Fix`) handle pg_query's convention of `StmtLen` excluding the trailing `;`.
## ✅ V3 — LSP + VSCode
- `pkg/lsp`: JSON-RPC 2.0 over stdio; `textDocument/formatting` (full document), `publishDiagnostics` on every open/change, `textDocument/codeAction` quick-fixes, lifecycle (initialize/shutdown/exit). No external deps.
- `cmd/pgtidy/lsp.go`: `pgtidy lsp` subcommand; config discovered from cwd.
- `editors/vscode/`: TS extension using `vscode-languageclient`; launches `pgtidy lsp` via stdio; `.pgsql` mapped to `sql` language; `pgtidy.path` / `pgtidy.enable` settings.
- _Range formatting: future._
- Full capability inventory, runtime-verified gaps, and ranked next steps in `docs/lsp-status.md` (issue #3).
## ✅ V4 — DataGrip
- `editors/datagrip/`: Gradle-based JetBrains plugin targeting DataGrip 2024.3+ via LSP4IJ.
- `build.gradle.kts` / `settings.gradle.kts` / `gradle.properties` — IntelliJ Platform Gradle Plugin v2.
- `plugin.xml` — registers `PgTidyServerFactory` as an LSP4IJ `<server>` extension and maps `*.sql`/`*.pgsql` to it.
- `PgTidyServerFactory.kt` + `PgTidyServerConnection.kt` — launches `pgtidy lsp` via `ProcessStreamConnectionProvider`.
- Requires LSP4IJ plugin installed in the IDE; `pgtidy` binary on PATH.
---
## ✅ Build / release (cross-cutting)
- `.goreleaser.yaml`: multi-platform matrix — linux/darwin × amd64/arm64 + windows/amd64; no CGO; ldflags version injection; draft GitHub release.
- `Makefile` extended: `snapshot` (local multi-platform build), `release` (publish), `vscode-compile`, `vscode-package`.
- `.github/workflows/ci.yml`: test + vet + gofmt check + goreleaser snapshot on every push/PR.
- `.github/workflows/release.yml`: goreleaser publish + VSCode `.vsix` artifact on `v*` tag.
- `make_release.sh` retained from boilerplate.
## Core invariants (must always hold — tested)
1. ✅ Lossless lex: `emit(Lex(src)) == src` (corpus round-trip).
2. ✅ Semantic equivalence: formatting changes only trivia/layout (corpus token-stream check).
3. ✅ Idempotence: `fmt(fmt(x)) == fmt(x)` (corpus + CLI tests).
4. ✅ Graceful degradation: unparsable spans pass through verbatim (Raw nodes + verbatim body).
## ✅ DML statement formatting (top-level)
- `pkg/format/dml.go`: `formatDML` formats top-level SELECT/INSERT/UPDATE/DELETE/WITH.
- Clause-per-line layout at depth-0 boundaries; handles: SELECT, FROM, WHERE, HAVING,
GROUP BY, ORDER BY, LIMIT, OFFSET, RETURNING, JOIN variants (LEFT/RIGHT/INNER/FULL/
CROSS/NATURAL, optional OUTER), ON CONFLICT, UNION/INTERSECT/EXCEPT, SET, VALUES,
INSERT INTO, DELETE FROM, WITH (including multi-CTE bodies).
- SELECT, SET, and RETURNING bodies formatted as leading-comma column lists.
- Keywords inside subqueries (paren depth > 0) cased correctly via `dmlInline`.
- `needSpace`: added `parenKws` set (AS, EXISTS, IN, NOT, LIKE, ILIKE, SIMILAR) so
these keywords get a space before `(` instead of the no-space function-call rule.
- Printer `writeItem`: detects DML-starting Raw nodes and routes to `formatDML`.
- Tests: 13 targeted golden tests + idempotence sweep in `pkg/format/dml_test.go`.
- Note: table name before `(` in INSERT column list is indistinguishable from a
function call at the token level — formatted without space (known limitation).
- Note: SQL keywords inside PL/pgSQL function bodies remain lowercase (matching
the corpus golden files); casing is applied only to top-level DML.
- _Still TODO: Wadler Doc-IR printer for width-aware wrapping of long lines._
- _Still TODO: LSP range formatting._
## ✅ Config expansion — DataGrip settings parity
Reference: `PostgresCodeStyleSettings` mapping in `docs/plan.md`.
### ✅ Extended `pkg/config` fields
Added to `Style` struct, `yamlFile`, and `Load()` in `pkg/config/config.go`:
- New types: `WrapMode` (`always`|`when_long`|`never`), `Placement` (`same_line`|`new_line`)
- **Casing**: `AliasCase`, `BuiltinCase`, `CustomTypeCase` — all default `lower`
- **Query layout**: `AlignColumns`, `AlignLineComments`, `SelectAlignAs`, `SetAlignEqual`,
`IndentJoin`, `JoinIndentSize`, `WhereWrap`, `WhereAndOrIndent`
- **Subqueries**: `SubqueryOpening`, `SubqueryContent`, `SubqueryClosing`, `SubquerySpaceBeforeParen`
- **INSERT**: `InsertCollapseValues`
- **Routines**: `AlignParamTypes`, `RoutineAsWrap`
- **PL/pgSQL**: `PlpgsqlMaxBlankLines`, `PlpgsqlDeclareAlignType`, `PlpgsqlDeclareAlignEq`,
`PlpgsqlIfThenNewline`, `PlpgsqlLoopCollapse`
- **Expressions**: `BinaryOpAlign`, `SpaceAfterCommaInCalls`, `CaseWhenWrap`, `CaseEnd`,
`CaseCollapse`, `RecordSpaceBeforeParen`
- `docs/config/default.pgtidy.yaml` updated with all new keys and comments.
### ✅ Casing engine — alias and built-in classification
`pkg/format/keywords.go`: added `builtinFunctions` set (COALESCE, MAX, MIN, NOW, …).
`pkg/format/format.go`: `caseTextCtx` uses context — `prev` token and `nextIsLParen` flag
to route ident tokens through `AliasCase` (after AS) or `BuiltinCase` (before `(`).
`inline()` and `dmlInline()` pass context to `caseTextCtx`.
### ✅ Formatter — query layout settings (`pkg/format/dml.go`)
- `indent_join` + `join_indent_size`: JOIN clause indented by `JoinIndentSize × Indent`.
- `where_wrap` + `where_and_or_indent`: `dmlWhereClause` splits AND/OR conditions; `always`
puts each condition on its own line indented under WHERE; `never` keeps inline.
- `set_align_equal`: `dmlColListSet` pads LHS of SET items so `=` signs align.
- `align_columns` + `select_align_as`: `dmlColListSelect` + `alignSelectItems` pads
SELECT expressions so AS keywords and aliases align vertically.
- `space_after_comma_in_calls` applied in `dmlInline`.
- `binary_op_align` registered in config (enforcement in WHERE/expression context deferred).
### ✅ Formatter — subquery formatting
`pkg/format/dml.go`: `dmlIsSubqueryOpen` detects a `(` immediately followed by `SELECT`/
`WITH` (derived tables, scalar subqueries, `IN`/`EXISTS`/`ARRAY(...)` subqueries — a plain
value tuple like `IN (1, 2, 3)` is left alone). `dmlInline` splices these in via
`dmlWrapSubquery`, which recursively formats the inner tokens with `formatDML` and wraps
them per `subquery_content`/`subquery_closing`; `dmlWriteSubquerySep` handles
`subquery_opening` (same_line/new_line) and additively applies `subquery_space_before_paren`
(only adds a space where one wouldn't already be there — never removes the space `IN`/
`EXISTS`/`AS` already get). `formatCTEDef` now calls the same `dmlWrapSubquery` helper
instead of a hardcoded new_line-only layout, so CTE bodies honor the config too (the
`AS (` space itself stays unconditional — that's fixed CTE syntax, not the subquery-space
setting). Nested subqueries-in-CASE and CASE-in-subqueries recurse correctly. Not
column/Doc-IR-aligned (documented "fixed-layout, not Wadler" limitation) — wrapped content
is indented one level relative to its own local frame, which composes correctly under
JOIN/WHERE/SELECT-list embedding but isn't perfectly column-aligned for deeply nested cases.
Tests: `TestDMLSubquery*`, `TestDMLCTEUsesSubqueryConfig` in `pkg/format/dml_test.go`.
### ✅ Formatter — INSERT VALUES collapse
`dmlValuesClause` (`pkg/format/dml.go`): when `insert_collapse_values` is `true` (default)
multi-row `VALUES` stays packed on one line (matches prior behavior); when `false`, each row
gets its own line via the shared `dmlCommaList` helper (same leading/trailing-comma layout
as SELECT/SET lists). Single-row VALUES is unaffected either way.
Tests: `TestDMLInsertValues*`.
### ✅ Formatter — routine param alignment (`pkg/format/format.go`)
- `align_param_types`: `alignParamTypes()` pads param names so type columns align; default `false` (house style: no type-column alignment in param lists).
- `routine_as_wrap`: when `false`, AS stays on the same line as the last option clause.
- Golden file `testdata/corpus/test_a.pgsql` updated to reflect aligned params.
### ✅ Formatter — PL/pgSQL body settings (`pkg/format/body.go`)
- `plpgsql_max_blank_lines`: blank-line runs capped at the configured limit; default `1`.
- `plpgsql_declare_align_type` + `plpgsql_declare_align_eq`: two-pass declare formatter
measures name/type widths then pads for alignment; `writeDeclareAligned` helper. Both
default `true` (house style). The `=` column is padded only to the widest type among
declarations that actually carry an assignment, so a lone `x text = '…';` stays tight.
- `plpgsql_if_then_newline`: when `false`, `joinThenToCondition` merges THEN onto the
preceding condition line.
- `plpgsql_loop_collapse`: `tryCollapseLoop` detects empty FOR/WHILE loop bodies and
collapses them to one line.
- CRLF normalization in trivia emission (comment text, body trivia before DECLARE).
### ✅ Formatter — expression settings (case_when_wrap, case_end, case_collapse, record_space_before_paren)
`pkg/format/dml.go`: `dmlInline` detects `CASE` tokens (`dmlIsCaseStart`/`dmlMatchCaseEnd`,
tracking nested-CASE depth so an inner `CASE…END`'s own `WHEN`/`THEN`/`ELSE` don't get
mistaken for the outer one's boundaries) and splices in `dmlFormatCase`. `dmlSplitCase`
breaks the body into operand/WHEN/THEN/ELSE segments (each rendered via a recursive
`dmlInline` call, so subqueries and nested CASEs inside a branch format correctly too).
Both the simple (`CASE x WHEN ...`) and searched (`CASE WHEN ...`) forms are supported.
- `case_when_wrap` (default `false`): `false` keeps everything on one line (unchanged
default behavior); `true` puts each `WHEN … THEN …` and `ELSE` on its own line, indented
one level.
- `case_end` (default `new_line`): placement of the closing `END` when wrapped —
`new_line` on its own line, `same_line` glued to the last WHEN/ELSE line.
- `case_collapse` (default `false`): when `true` *and* the fully-inlined rendering is
≤ `caseCollapseWidth` (60 chars), keeps the wrapped CASE on one line anyway, overriding
`case_when_wrap`; longer CASEs still wrap. (No Doc-IR/line-width awareness exists yet, so
this is a fixed length threshold rather than a true "does it fit the line" check.)
- `record_space_before_paren` (default `false`): scoped to the `ROW` keyword specifically
(`ROW(1, 2)` vs `ROW (1, 2)`) — bare `(a, b)` record literals are indistinguishable from
grouping parens at the token level, so this setting only fires on an explicit `ROW(`.
Tests: `TestDMLCase*`, `TestDMLRecordSpaceBeforeParen`.
### ⬜ DataGrip XML import/export (optional, V4+)
`pgtidy config import --datagrip <settings.xml>` / `pgtidy config export --datagrip`
not implemented.
---
## Open risks
- `go-pgquery` tracks PG17 (not PG18) — fine for lint; irrelevant to formatter path.
- Leading-comma + one-per-line is a first-class style option, not an afterthought.