equals the parent directory name.
- **Plain-text output**: The optional rendered output (HTML, printable
PDF, or other) must use print-friendly typography. No emojis appear in
any rendering template
handles the task better than a custom solution — document-creation skills (`docx`, `xlsx`, `pptx`, `pdf`), visualisation tools, code execution, maps/places, frontend design guidance, and so on. If the task
line about the step before, no history), `RenderedCV`
(`tex`, `pdf_base64`, `document`), `CoverLetter`, `ServerStatus`. The CV
document is a raw flat dict with no schema — what the model wrote with
exceeding limits before processing
- **Content-Type Validation**:
- Validate file extension against allowlist (e.g., `.jpg`, `.pdf`, `.docx`)
- Check magic bytes match expected file type (first few bytes of file)
- Use specialized
variables for theme customization (see `slides/themes/custom-default.css`).
- Ensure compatibility with Marp CLI v4 for PDF and presentation generation
Authentication**: Use `@permission_classes([IsAuthenticated])` for protected endpoints
- **File Uploads**: Use `MultiPartParser, FormParser` for PDF uploads
- **OpenAI Key**: Always use `os.getenv("OPENAI_API_KEY")` from backend environment, never user-provided
labels.
- Include a legend when more than one dataset is plotted.
- Save figures as **PDF** for vector graphics (publications) and PNG
for quick inspection. Do both when in doubt.
- Always
operator-configured TSA / OCSP / CRL endpoints,
reproducible output, deterministic releases). Faithful thin wrapper: every PDF feature lives in
the engine; never reimplement one, never over-promise (state engine limits
when adding a new MCP server or debugging "no output" / early-exit failures. |
| pdf | `**/pdf/**/*.pdf, **/reports/**/*.pdf, **/pdf*.py, **/pdf*.ts, **/pdf*.js, **/generate*pdf*, **/*pdf*generator*` — Use this
multimodal RAG (Retrieval-Augmented Generation) system** that enables intelligent search across documents (PDF, DOCX), images, and audio files. The architecture follows a **local-first approach** - all processing happens on-device
Python Scripts**:
- Use `scripts/hl_v3_final/hl_lib.py` for all sentence locating, bounding rect generation, and tri-modal PDF annotations.
- Use `scripts/provider_llm.py` and `scripts/provider_vision.py` for model provider abstraction.
- Verify tests with `python3 scripts/hl_v3_final/test_hl_lib.py
titles deck.pptx # read the assertions as prose: do they argue?
soffice --headless --convert-to pdf deck.pptx && pdftoppm -png -r 70 deck.pdf p
```
The rendered output is the truth
file size, and per-page dimensions as JSON; no comparison. Optional `password`.
Supported formats: PDF, DOCX, XLSX, PPTX, ODT, ODS, ODP, RTF, TXT, HTML, and 30+ more.
## Building this repo
package `src/acikpoz/`. The pipeline is layered so parsing logic is testable without a real PDF:
- `model.py` — `Poz` dataclass; every possibly-missing field is `Optional`, defaulting to `None`.
- `parser.py` — coordinate/geometry parser