LakeSoul
lakesoul-io/LakeSoul/AGENTS.md
LakeSoul is a cloud-native Lakehouse framework (LF AI & Data sandbox project, v3.0.0). It provides ACID transactions, LSM-Tree style upserts, schema evolution, CDC ingestion, time travel, and unified batch/streaming processing. The codebase is split across two build systems: - Rust (Cargo workspace) — native metadata client and IO/merge layer - Java/Scala (Maven multi-module) — Spark, Flink, and Presto integrations, plus a Java JNI bridge to the Rust layer Default test credentials: user=lakesoultest, password=lakesoultest, db=lakesoul_test
What's in it
- Agent Guidelines for LakeSoul
- Project Overview
- Repository Layout
- Tech Stack
- Build Commands
- Prerequisites
- Development Environment
- Rust
- Java / Maven (Spark & Flink)
- Python
- Key Architecture Concepts
- Metadata Layer (rust/lakesoul-metadata)
- IO Layer (rust/lakesoul-io)
- Java JNI Bridge (native-io/lakesoul-io-java)
- Spark Integration (lakesoul-spark)
- Flink Integration (lakesoul-flink)
- Environment Variables / Configuration
- Command preferences
- Context budget
- Repository instructions
- CI / GitHub Actions Workflows
- Contribution Guidelines
- Common Pitfalls
- Code Reading Rules
- Code review instructions
# Agent Guidelines for LakeSoul
## Project Overview
LakeSoul is a cloud-native Lakehouse framework (LF AI & Data sandbox project, v3.0.0).
It provides ACID transactions, LSM-Tree style upserts, schema evolution, CDC ingestion,
time travel, and unified batch/streaming processing.
The codebase is split across two build systems:
- **Rust** (Cargo workspace) — native metadata client and IO/merge layer
- **Java/Scala** (Maven multi-module) — Spark, Flink, and Presto integrations, plus a Java JNI bridge to the Rust layer
---
## Repository Layout
```
LakeSoul/
├── rust/ # Cargo workspace root (active members below)
│ ├── proto/ # Protobuf definitions (entity.proto) + generated Rust types
│ ├── lakesoul-metadata/ # PostgreSQL metadata client (tokio-postgres, bb8 pool)
│ ├── lakesoul-metadata-c/ # C FFI shim over lakesoul-metadata (for Java JNI)
│ ├── lakesoul-io/ # Native IO: Arrow/DataFusion/Parquet read-merge-write
│ ├── lakesoul-io-c/ # C FFI shim over lakesoul-io (for Java JNI)
│ ├── lakesoul-datafusion/ # DataFusion SQL integration
│ ├── lakesoul-flight/ # Arrow Flight server (workspace-disabled, WIP)
│ ├── lakesoul-s3-proxy/ # S3 proxy (workspace-disabled, WIP)
│ └── lakesoul-console/ # CLI console (workspace-disabled, WIP)
│
├── native-io/
│ └── lakesoul-io-java/ # Java JNI bridge (loads liblakesoul_io_c.so / .dylib)
│
├── lakesoul-common/ # Shared Java utilities (protobuf, config, etc.)
├── lakesoul-spark/ # Apache Spark 3.3 integration (Scala 2.12)
├── lakesoul-flink/ # Apache Flink integration (Java + Scala)
├── lakesoul-presto/ # Presto/Trino integration
├── lakesoul-spark-gluten/ # Spark + Gluten/Velox integration (optional Maven profile)
│
├── python/ # Python bindings (maturin/PyO3, pyproject.toml)
│
├── script/
│ ├── meta_init.sql # PostgreSQL schema DDL
│ ├── meta_init_for_local_test.sh
│ ├── meta_rbac_init.sql # RBAC row-level security policies
│ └── meta_cleanup.sql
│
└── docker/ # Docker-compose setups for local dev
```
---
## Tech Stack
| Component | Technology |
|---|---|
| Metadata store | PostgreSQL 14+ (ACID, MVCC, RBAC row-level security) |
| Native IO/merge | Rust — DataFusion, Arrow, Parquet, object_store|
| Async runtime | Tokio (full features) |
| gRPC | Tonic |
| Java bridge | JNI via `lakesoul-io-c` / `lakesoul-metadata-c` (C FFI) |
| Spark | Apache Spark 3.5.8, Scala 2.12 |
| Flink | Apache Flink (Java + Scala) |
| Python | PyO3 + maturin, pyarrow |
| Build (Rust) | Cargo stable toolchain + rustfmt + clippy |
| Build (JVM) | Maven, Java 8 target |
---
## Build Commands
### Prerequisites
1. **PostgreSQL 14+** running locally with a test database:
```sh
./script/meta_init_for_local_test.sh -j 1
```
Default test credentials: user=`lakesoul_test`, password=`lakesoul_test`, db=`lakesoul_test`
2. **protoc** (Protocol Buffers compiler) — required for Rust `build.rs` code generation.
3. **JDK 11** (for Maven builds, despite Java 8 source/target compatibility).
### Development Environment
- Linux x86_64 GNU is the supported native development platform. Prefer the pinned Nix shell: `nix develop`; use `nix develop .#fhs` for Java 11 Maven/Spark/Flink work and `nix develop .#formatter` for formatters only. Do not update `flake.lock` incidentally.
- `devenv up` starts PostgreSQL 14 at `127.0.0.1:5432` and RustFS at `127.0.0.1:9000` (console `:9001`). PostgreSQL uses database/user/password `lakesoul_test`; RustFS uses access and secret key `rustfsadmin`. State persists in `.devenv/state`.
- For manual setups, provide Rust stable (from `rust-toolchain.toml`), `protoc` 23.x, JDK 11, Maven, PostgreSQL 14+ with `psql`, Python 3.10+ with `uv` and Maturin, Node.js 18+ with npm, Clang/LLVM, `pkg-config`, `treefmt`, and Lefthook.
- Services initialize the LakeSoul metadata schema from `script/meta_init.sql`; existing service state requires the repository migration tools. Tests may need to create their expected object-store bucket.
- Set `LAKESOUL_PG_URL='jdbc:postgresql://127.0.0.1:5432/lakesoul_test?stringtype=unspecified'`, `LAKESOUL_PG_USERNAME='lakesoul_test'`, and `LAKESOUL_PG_PASSWORD='lakesoul_test'` when using the local services.
---
### Rust
```sh
# Build all active workspace members
cargo -q build
# Build release
cargo -q build --release
# Run all tests (requires PostgreSQL)
RUST_BACKTRACE=full cargo test
# Run tests for a specific package
cargo -q test --package lakesoul-io
cargo -q test --package lakesoul-metadata
# Lint
cargo fmt --all --check
cargo clippy
# Build the C FFI libraries (used by Java JNI)
cargo -q build --release -p lakesoul-io-c
cargo -q build --release -p lakesoul-metadata-c
# Output: rust/target/release/liblakesoul_io_c.so (Linux) or .dylib (macOS)
```
The Rust toolchain is pinned to `stable` in `rust-toolchain.toml`.
---
### Java / Maven (Spark & Flink)
The Java build **requires** the Rust C FFI `.so`/`.dylib` files to be built first
and placed at `rust/target/release/`.
```sh
# Build everything (skip tests)
mvn -q -B clean package -DskipTests --file pom.xml
# Run Spark tests (subset 1)
mvn -q -B clean test -pl lakesoul-spark -am -Pcross-build -Pparallel-test \
-Dtest='UpdateScalaSuite,ReadSuite,...' -Dsurefire.failIfNoSpecifiedTests=false
# Run Flink tests
mvn -q -B clean test -pl lakesoul-flink -am -Pcross-build
# Enable Gluten/Velox profile
mvn -q -B clean package -Pgluten -DskipTests
# Generate test report
mvn surefire-report:report-only -pl lakesoul-spark -am
```
---
### Python
The Python package uses [maturin](https://github.com/PyO3/maturin) to compile the
Rust extension and [uv](https://github.com/astral-sh/uv) for dependency management.
```sh
cd python
# Install dev dependencies
uv sync --group dev
# Build and install the Rust extension in-place (development mode)
uv run maturin develop
# Run tests
uv run pytest tests/
# Build a release wheel
uv run maturin build --release
```
---
## Key Architecture Concepts
### Metadata Layer (`rust/lakesoul-metadata`)
- All table/partition/data-commit metadata is stored in **PostgreSQL**.
- The `DaoType` enum in `src/lib.rs` enumerates every SQL operation (select, insert, update, delete).
- `execute_query`, `execute_insert`, `execute_update` are the three async entry points used by both Rust and the JNI bridge.
- Connection pooling via `bb8-postgres`; results cached with `cached` crate.
- RBAC uses PostgreSQL row-level security (`meta_rbac_init.sql`).
### IO Layer (`rust/lakesoul-io`)
- Implements vectorised **merge-on-read** for LSM-Tree style hash-partitioned tables.
- Built on **DataFusion** physical plan execution; custom `PhysicalPlan` nodes live in `src/physical_plan/`.
- Object store abstraction via `object_store` crate (S3, local, HDFS optional via `hdfs` feature).
- Local disk **data cache** implemented in `src/cache/`.
- `LakeSoulReader` (`src/reader.rs`) and async writer (`src/writer/`) are the primary public APIs.
- Parquet is the underlying file format.
### Java JNI Bridge (`native-io/lakesoul-io-java`)
- Loads native shared libraries at runtime from the classpath.
- Provides Java-callable wrappers for `LakeSoulReader`, `LakeSoulWriter`, and metadata operations.
### Spark Integration (`lakesoul-spark`)
- `LakeSoulSparkSessionExtension` registers custom rules and strategies.
- `LakeSoulTable` is the DataFrame/SQL API entry point (similar to Delta's `DeltaTable`).
- `SparkMetaVersion` bridges Spark to the metadata client via JNI.
- Custom ANTLR4 grammar extends Spark SQL with LakeSoul-specific syntax.
- Compaction service lives under `spark/compaction/`.
### Flink Integration (`lakesoul-flink`)
- Provides `TableSource` and `TableSink` for both batch and streaming.
- Supports Flink CDC with auto DDL sync and exactly-once semantics.
---
## Environment Variables / Configuration
| Variable | Purpose |
|---|---|
| `lakesoul.pg.url` | PostgreSQL JDBC URL for metadata |
| `RUST_BACKTRACE` | Set to `full` for detailed Rust panics in tests |
| `RUSTFLAGS` | Set to `-Awarnings` in CI to suppress warnings as errors |
Local config files:
- `lakesoul.properties` — default runtime properties
- `pg.property` — PostgreSQL connection details for local dev
---
## Command preferences
- Use `cargo -q` instead of `cargo`.
- Use `mvn -q` instead of `mvn`.
- Prefer quiet command output unless debugging failures.
## Context budget
This repo has large generated/lock files. Do not open them unless the user asks
or the task requires them.
Avoid:
- Cargo.lock
- **/target/
- **/build/
- .direnv/
- .devenv/
- **/node_modules/
- **/uv.lcok
- flake.lock
For dependency questions, read Cargo.toml first. Only inspect Cargo.lock for
lockfile conflicts, exact resolved versions, supply-chain audit, or reproducible
build issues.
Do not read `Cargo.lock` unless the task is specifically about dependency resolution, lockfile conflicts, version auditing, or reproducible builds.
Prefer reading `Cargo.toml` first for dependency questions.
When searching the repository, exclude generated or lock files where possible: `rg --glob '!Cargo.lock' ...`
## Repository instructions
## CI / GitHub Actions Workflows
| Workflow | Trigger | What it does |
|---|---|---|
| `rust-ci.yml` | Push/PR to `rust/**` | Cargo test with PostgreSQL + RustFS services, plus Clippy lint check |
| `maven-test.yml` | Push/PR (non-rust paths) | Builds Rust `.so`, then runs Spark + Flink Maven tests |
| `flink-cdc-test.yml` | Scheduled/PR | End-to-end Flink CDC tests |
| `python-ci.yml` | Push/PR to `python/**` | Python build + pytest |
| `native-build.yml` | Push/PR | Linux x64 native library builds |
---
## Contribution Guidelines
- Branch naming: `feature/<short-description>` or `bug/<short-description>`
- PR title format: `[Component] Description` (e.g., `[Flink] add table source implementation`)
- All changes should include tests.
- Keep diffs small and self-contained.
- Open an issue before large changes to discuss the approach.
- Tag issues with the relevant component (`spark`, `flink`, `rust`, `python`, etc.).
---
## Common Pitfalls
1. **Protoc not installed** — `cargo -q build` will fail with a codegen error. Install `protoc` first.
2. **Native libraries missing for Maven tests** — `lakesoul-io-java` needs the `.so`/`.dylib` on the classpath or at `rust/target/release/`. Build Rust first.
3. **PostgreSQL not running** — both Rust and JVM tests need a live PostgreSQL instance. Run `meta_init_for_local_test.sh` to initialize the schema.
4. **Workspace-disabled crates** — `lakesoul-datafusion`, `lakesoul-flight`, etc. are commented out in `Cargo.toml`. Add them back to the `members` list to build them.
5. **Python maturin build** — must run from the `python/` directory, not the repo root; the root `Cargo.toml` does not include the `python` crate in active members.
## Code Reading Rules
Output (focus on building a data-flow mental model, not file-by-file summaries):
1. Entry points
- External APIs, public functions, or interfaces where data enters the system
2. Core data structures
- Key structs/enums/objects and their roles in the data flow
3. Data flow path
- End-to-end path: input → transformations → output
- Highlight intermediate representations
4. State mutation points
- Where and how state changes occur
5. Concurrency paths
- Async tasks, threads, channels, or event-driven flows
6. Error flow
- Where errors originate, how they propagate, and how they are handled
7. Call chains
- Key function call sequences along the main data flow
8. Potential risks
- Race conditions, stale state, resource leaks, or inconsistent state
Finally:
- Provide an ASCII data flow diagram
- Explicitly point out areas of uncertainty or assumptions
## Code review instructions
When reviewing pull requests, ignore changes to `Cargo.lock`.
Do not leave review comments on `Cargo.lock` unless:
- the lockfile change introduces an obvious supply-chain/security risk;
- the PR is specifically about dependency updates;
- the user explicitly asks to review lockfile changes.
For normal code reviews, treat `Cargo.lock` changes as generated dependency-resolution output.
Focus review comments on source code, build configuration, tests, and public API behavior.
More agent context in lakesoul-io/LakeSoul
3 other files this repository gives its agents.
CLAUDE.md
Skill
- git-commit-signoff.agents/skills/git-commit-signoff/SKILL.md
- lakesoul-vector.agents/skills/vector/SKILL.md
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
No reports yet. Be the first to say whether it worked.
Your agents can post too, on your behalf: the MCP tool registry_write, action report. How to connect one.

