agents / data-extraction-agent
Natsummerance/agents/data-extraction-agent/AGENTS.md
AGENTS.md3 starsChanged 38 days ago
# AGENTS.md - 工作流程和场景定义 ## 🚀 启动流程 1. 读取 IDENTITY.md 2. 读取 SOUL.md 3. 不需要询问许可,直接执行 --- ## 📋 核心工作流 ### 场景 1:PDF 结构化提取 读取 ODL 工具帮助 → 构建最小命令 → 执行提取 → 验证结果完整性 → 输出 JSON/Markdown。 ### 场景 2:LLM 友好网页爬取 配置 Crawl4AI → 异步抓取目标页面 → 清洗为 LLM 可用的 Markdown/结构化数据。 ### 场景 3:反反爬隐身抓取 使用 Scrapling 绕过 Cloudflare Turnstile 等防护 → 隐身浏览模式 → 提取目标数据。 --- ## 🛑 错误处理 | 状况 | 行为 | |------|------| | PDF 加密或损坏 | 报告原因,建议替代方案 | | 目标网站反爬过于严格 | 说明限制,不尝试非法绕过 |
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

