# Development specification: AI Content Pipeline v1 ## 1. Objective Build an extendable content production pipeline for writers and editors. The system accepts a short article description, asks boundary-setting questions, generates an article plan, stops for editor review, then continues through research, evidence gathering, section scaffolding, visual generation, SEO, linguistic review, final approval, and Git-backed publishing into a Next.js content repository. v1 is an internal or single-tenant editorial tool. It is not a customer-facing multi-tenant SaaS product. Chosen architecture: ```text React / Next.js UI FastAPI backend LangGraph workflow engine Postgres state store Agent Runner Service using Codex CLI subscription authentication Optional Claude Code runner for internal-only workflows Object storage for files and assets Git-backed publishing adapter layer for multiple target websites ``` LangGraph is appropriate because it supports durable execution, human-in-the-loop workflow control, and stateful orchestration. These are central to the required “stop, wait for edits, resume” behavior. ([docs.langchain.com][1]) Codex should be the primary subscription-based agent runner because Codex supports ChatGPT sign-in for subscription access, while Codex CLI also supports ChatGPT account login, API key login, or access-token login. For this project, use one managed runner identity per environment through ChatGPT subscription login or enterprise access-token auth, not usage-based API-key calls. ([OpenAI Разработчики][2]) Claude Code can be supported as an optional internal runner, but not as the main backend for routing third-party user jobs through consumer subscription credentials. Anthropic documents OAuth and API-key authentication as serving different purposes and states limits around third-party product/service usage. ([Claude Code][3]) --- # 2. Product scope ## In scope for v1 ```text 1. Article intake 2. Boundary question generation 3. Plan generation 4. Plan approval gate 5. Web research task 6. Evidence matrix generation 7. Parallel section scaffolding 8. Hero image prompt generation 9. Table and diagram specification generation 10. SEO pass 11. Linguistic and tone review 12. Final approval gate 13. Git-backed publish commit creation 14. Multi-site configuration 15. Job history and audit trail 16. Codex CLI-based agent execution ``` ## Out of scope for v1 ```text 1. Fully autonomous publishing without human approval 2. Multi-language localization 3. A/B testing 4. Content performance analytics 5. Legal/compliance review workflows 6. Real-time collaborative editing 7. Direct Anthropic/OpenAI API-token usage for generation 8. Complex media editing pipeline 9. Owning CI/CD deployment orchestration 10. Multi-tenant customer isolation 11. Per-user Codex runner credentials ``` The publishing step in v1 should create a **direct commit to the configured production branch** of a Git-backed Next.js site repository after final editor approval and a best-effort content-shape dry run. The target repository's existing CI/CD owns deployment. --- # 3. Core user roles ## Admin Can configure target websites, pipeline parameters, agent runner profiles, prompt versions, publishing YAML, and site-specific transformation/upload scripts. ## Editor Can create briefs, answer boundary questions, run the pipeline, edit intermediate results, approve plans, edit drafts, approve final content, and create publish commits. Only Admins can edit upload scripts and pipeline configuration. Admins are fully trusted code operators because admin-defined transformation scripts run directly on the runner host inside checked-out site repositories. --- # 4. High-level workflow ```text ARTICLE_BRIEF_CREATED | v BOUNDARY_QUESTIONS_GENERATED | v BOUNDARY_ANSWERS_SUBMITTED | v PLAN_GENERATED | v PLAN_REVIEW_REQUIRED | +--> editor requests changes --> PLAN_REVISION_REQUIRED --> PLAN_GENERATED | +--> editor approves v RESEARCH_RUNNING | v EVIDENCE_MATRIX_READY | v PARALLEL_PRODUCTION_RUNNING | +--> section scaffolds +--> hero image prompt +--> table specs +--> diagram specs +--> SEO brief +--> FAQ block | v DRAFT_ASSEMBLED | v SEO_AND_LANGUAGE_REVIEW_READY | v FINAL_REVIEW_REQUIRED | +--> editor requests changes --> FINAL_REVISION_REQUIRED | +--> editor approves v PUBLISH_DRY_RUN_REQUIRED | v PUBLISH_COMMIT_READY | v PUBLISH_COMMIT_CREATED | v DONE ``` --- # 5. System architecture ```text Frontend React / Next.js | v Backend API FastAPI | +--> Auth service +--> Article service +--> Site config service +--> Workflow service +--> Review service +--> Asset service +--> Publishing service | v Workflow engine LangGraph | v Queue Redis Queue / Celery / Dramatiq | v Agent Runner Service | +--> Codex CLI runner +--> Optional Claude Code runner +--> Web research script +--> Diagram script +--> Git-backed publishing script | v Storage +--> Postgres +--> Object storage +--> Optional vector store ``` FastAPI’s built-in background tasks are useful for small post-response actions, but this system should use a real queue for long-running agent jobs, because article generation, research, retries, and approval waits are workflow-level tasks rather than short request-level tasks. FastAPI supports background tasks, but the durable workflow should live in LangGraph plus a queue. ([fastapi.tiangolo.com][4]) --- # 6. Main components ## 6.1 Frontend UI ### Required screens ```text 1. Dashboard 2. New article brief 3. Boundary questions 4. Plan review 5. Research pack view 6. Evidence matrix view 7. Draft editor 8. Asset review 9. SEO review 10. Final approval 11. Publishing settings 12. Site configuration 13. Workflow history 14. Media library 15. Generic Markdown/MDX rich preview ``` v1 frontend should feel like a full editorial workspace, including rich preview, media management, and advanced draft editing. If timeline forces scope reduction, scheduling slips before rich preview. ### Dashboard fields ```text Article title Target website Workflow status Assigned editor Last updated Next required action Publishing status ``` ### Plan review UI must support ```text Approve plan Request revision Edit plan directly Add section-level notes Add source requirements Add excluded sources Add visual requirements Change tone Change target audience Change SEO keyword ``` ### Final review UI must support ```text Approve final draft Request revision Edit title Edit meta description Edit article body Edit images Edit tables Edit diagrams Edit internal links Edit frontmatter, content path, category, and tags Create publish commit ``` --- ## 6.2 Backend API Backend stack: ```text FastAPI Postgres SQLAlchemy or SQLModel Pydantic models Redis Queue / Celery / Dramatiq LangGraph workflow module Object storage SDK Git publishing adapter interfaces ``` ### API design principles ```text 1. Every workflow-changing action must be explicit. 2. Every human approval must be recorded. 3. Every agent output must be validated before entering the article state. 4. Every generated claim must be traceable to evidence or marked unsupported. 5. Every target website must be driven by config, not hardcoded logic. 6. Article/domain tables are authoritative for product state; LangGraph checkpoints are execution state only. 7. Every approval-relevant plan or draft edit creates a new immutable version. ``` --- # 7. Data model ## 7.1 Article ```json { "id": "uuid", "target_site_id": "uuid", "created_by": "uuid", "assigned_editor_id": "uuid", "status": "PLAN_REVIEW_REQUIRED", "brief_description": "string", "working_title": "string", "language": "en", "content_type": "longform_guide", "primary_keyword": "string", "created_at": "datetime", "updated_at": "datetime" } ``` ## 7.2 TargetSite ```json { "id": "uuid", "name": "B2B SaaS Blog", "slug": "b2b_saas_blog", "publishing_type": "git_next", "default_language": "en", "brand_voice": "direct, useful, evidence-backed", "audience": "B2B SaaS founders and content leads", "seo_rules": {}, "visual_rules": {}, "source_rules": {}, "publishing_rules": { "repository_url": "git@github.com:example/site.git", "production_branch": "main", "content_format": "mdx", "content_path_template": "content/articles/{slug}.mdx", "asset_path_template": "public/articles/{slug}/{filename}", "frontmatter_mapping": {}, "transform_script_version_id": "uuid" }, "created_at": "datetime", "updated_at": "datetime" } ``` ## 7.3 BoundaryQuestion ```json { "id": "uuid", "article_id": "uuid", "question": "Who is the article for?", "question_type": "audience", "answer": "Content operations leads", "required": true, "sort_order": 1 } ``` ## 7.4 ArticlePlan ```json { "id": "uuid", "article_id": "uuid", "version": 1, "status": "PENDING_REVIEW", "title_options": [], "recommended_title": "string", "search_intent": "string", "thesis": "string", "sections": [], "claims_to_prove": [], "visuals_needed": [], "seo_notes": {}, "risks": [], "editor_notes": [], "created_at": "datetime" } ``` ## 7.5 PlanSection ```json { "id": "uuid", "article_plan_id": "uuid", "sort_order": 1, "heading": "string", "purpose": "string", "key_points": [], "claims_to_support": [], "evidence_needed": [], "visuals_needed": [], "target_word_count": 500 } ``` ## 7.6 EvidenceItem ```json { "id": "uuid", "article_id": "uuid", "source_title": "string", "source_url": "string", "source_domain": "string", "source_type": "official_docs | research | news | company_blog | other", "source_quality_score": 0.9, "summary": "string", "relevant_quotes": [], "supports_claims": [], "used_in_sections": [], "artifact_manifest_id": "uuid", "snapshot_object_key": "s3://bucket/research-runs/run_123/source_001.json", "snapshot_content_hash": "sha256", "retrieved_at": "datetime" } ``` ## 7.7 Claim ```json { "id": "uuid", "article_id": "uuid", "section_id": "uuid", "claim_text": "string", "support_status": "SUPPORTED | UNSUPPORTED | NEEDS_REVIEW", "evidence_item_ids": [], "risk_level": "low | medium | high" } ``` ## 7.8 ArticleDraft ```json { "id": "uuid", "article_id": "uuid", "version": 1, "title": "string", "slug": "string", "meta_title": "string", "meta_description": "string", "body_markdown": "string", "faq_block": [], "schema_json": {}, "internal_links": [], "external_links": [], "status": "DRAFT_ASSEMBLED" } ``` ## 7.9 Asset ```json { "id": "uuid", "article_id": "uuid", "asset_type": "hero_image | diagram | table | inline_image", "title": "string", "prompt": "string", "file_url": "string", "alt_text": "string", "caption": "string", "status": "PENDING | GENERATED | APPROVED | REJECTED" } ``` ## 7.10 WorkflowEvent ```json { "id": "uuid", "article_id": "uuid", "event_type": "PLAN_APPROVED", "actor_type": "user | system | agent", "actor_id": "uuid", "payload": {}, "created_at": "datetime" } ``` ## 7.11 AgentJob ```json { "id": "uuid", "article_id": "uuid", "job_type": "PLAN_GENERATION", "agent_profile": "codex_subscription_default", "status": "QUEUED | RUNNING | SUCCEEDED | FAILED | CANCELLED", "workspace_path": "string", "input_files": [], "output_files": [], "error_message": "string", "started_at": "datetime", "finished_at": "datetime" } ``` ## 7.12 ResearchRunManifest ```json { "id": "uuid", "article_id": "uuid", "agent_job_id": "uuid", "s3_prefix": "research-runs/article_123/run_001/", "artifacts": [ { "artifact_type": "source_section_snapshot", "source_url": "https://example.com/article", "object_key": "research-runs/article_123/run_001/source_001.json", "content_hash": "sha256", "metadata": { "search_path": [], "agent_selection_criteria": "string", "head_metadata": {}, "server_ip_address": "string", "domain_whois_owner": "string" } } ], "created_at": "datetime" } ``` Research source sections and metadata are stored indefinitely in object storage, not as large database blobs. Postgres stores manifests, object keys, hashes, source URLs, and artifact types. ## 7.13 PublishCommit ```json { "id": "uuid", "article_id": "uuid", "target_site_id": "uuid", "repository_url": "string", "branch": "main", "commit_sha": "string", "content_bundle_manifest": {}, "status": "PUBLISH_COMMIT_CREATED | PUBLISH_VERIFICATION_FAILED", "deployment_status": "UNKNOWN | DISCOVERED | FAILED | SUCCEEDED", "created_at": "datetime" } ``` ## 7.14 ScriptConfigVersion ```json { "id": "uuid", "target_site_id": "uuid", "version": 1, "author_id": "uuid", "yaml_config": {}, "transform_script": "string", "diff_from_previous": "string", "rollback_target_version_id": "uuid", "created_at": "datetime", "activated_at": "datetime" } ``` Every Admin script/config change stores author, timestamp, diff, and rollback target. v1 does not require second Admin approval for activation. --- # 8. Agent Runner Service ## 8.1 Purpose Run subscription-authenticated local agents in isolated workspaces without exposing direct model API keys to the application. Primary runner: ```text Codex CLI with one managed ChatGPT subscription or enterprise access-token identity per environment ``` Optional runner: ```text Claude Code for internal-only workflows ``` ## 8.2 Runner requirements ```text 1. Create isolated workspace per job. 2. Write input files into workspace. 3. Run allowed CLI command. 4. Capture stdout, stderr, exit code. 5. Enforce timeout. 6. Validate output JSON against schema. 7. Store generated artifacts. 8. Return result to backend. 9. Record full job log. 10. Prevent access to unrelated article workspaces. ``` ## 8.3 Example workspace ```text /workspaces/article_123/job_plan_generation_001/ input/ brief.json target_site.json boundary_answers.json output_schema.json instructions.md output/ plan.json logs/ stdout.log stderr.log command.json ``` ## 8.4 Example Codex job command ```bash codex exec "Read input/instructions.md. Use input files. Produce output/plan.json only. Follow input/output_schema.json exactly." ``` ## 8.5 Agent output validation All agent outputs must pass validation before workflow state changes. Validation rules: ```text 1. JSON must be valid. 2. JSON must match expected schema. 3. Required fields must be present. 4. No empty section plan. 5. Every generated claim must be tied to a plan section. 6. Research output must include source URL, retrieval timestamp, artifact manifest reference, and content hash. 7. Draft cannot include unsupported factual claims unless marked for editor review. 8. Final approval cannot proceed while high-risk unsupported claims remain unresolved. ``` --- # 9. Workflow nodes ## 9.1 Intake node Input: ```json { "brief_description": "string", "target_site_id": "uuid", "content_type": "longform_guide", "primary_keyword": "optional string" } ``` Output: ```text Article created Initial workflow event stored Status: ARTICLE_BRIEF_CREATED ``` Acceptance criteria: ```text Given a valid brief When the user submits it Then an article record is created And workflow status is ARTICLE_BRIEF_CREATED And target site config is attached ``` --- ## 9.2 Boundary question node Purpose: Generate 5-10 questions to define scope, audience, angle, constraints, SEO, and evidence expectations. Required question categories: ```text Audience Purpose Reader outcome Depth Tone Excluded topics Primary keyword Competitor angle Evidence standard Visual expectations ``` Output: ```json { "questions": [ { "question": "Who is the article for?", "question_type": "audience", "required": true } ] } ``` Acceptance criteria: ```text Questions are generated from brief plus target site config. Questions are editable before submission. Required questions block plan generation until answered. ``` --- ## 9.3 Plan generation node Purpose: Generate article production contract. Plan must include: ```text Title options Recommended title Reader persona Search intent Thesis Detailed section outline Claims to prove Evidence needs Visual needs SEO notes Tone guidance Risks Suggested internal links ``` Output schema: ```json { "title_options": ["string"], "recommended_title": "string", "reader_persona": "string", "search_intent": "informational | commercial | navigational | transactional", "thesis": "string", "sections": [ { "heading": "string", "purpose": "string", "key_points": ["string"], "claims_to_support": ["string"], "evidence_needed": ["string"], "visuals_needed": ["string"], "target_word_count": 500 } ], "seo": { "primary_keyword": "string", "secondary_keywords": ["string"], "meta_title_draft": "string", "meta_description_draft": "string" }, "risks": ["string"] } ``` Acceptance criteria: ```text Plan contains at least 4 sections. Each section has purpose, key points, and evidence needs. Plan is saved as version 1. Workflow stops at PLAN_REVIEW_REQUIRED. No research starts before approval. ``` --- ## 9.4 Plan review gate Purpose: Human approval checkpoint. Allowed editor actions: ```text Approve Request revision Edit directly Reject article Add notes Change target site Change content type ``` Acceptance criteria: ```text Workflow cannot continue until plan is approved. Every approval or revision request is stored as WorkflowEvent. Plan revisions create new versions. Previous versions remain accessible. ``` --- ## 9.5 Research node Purpose: Gather information from the web after plan approval through an explicit web search/fetch script, not free-form agent browsing alone. Inputs: ```text Approved plan Target site source rules Claims to prove Required source types Excluded domains Editor notes ``` Output: ```text Research pack Evidence matrix Source summaries Claim-source mapping Unsupported claims list Research run manifest with S3 object keys and content hashes ``` Research artifact storage: ```text 1. Store broader discovered source content sections, not only final-used evidence snippets. 2. Store page metadata, search path to the article, agent criteria used to select it as evidence, page metadata, server IP address, and domain WHOIS owner. 3. Store source sections and metadata indefinitely in object storage. 4. Store only manifest records, S3 object keys, source URLs, content hashes, and artifact types in Postgres. ``` Evidence quality scoring: ```text Official documentation: high Peer-reviewed research: high Government or standards body: high Reputable news/reporting: medium-high Company blog: medium Generic SEO blog: low-medium Forum/reddit: context only Unattributed content: low ``` Acceptance criteria: ```text Each factual claim has at least one source or is marked unsupported. Each source has URL, title, domain, source type, summary, and retrieval timestamp. Each research run writes an object-storage manifest. If enough acceptable evidence cannot be found for the approved plan, workflow returns to PLAN_REVISION_REQUIRED before drafting. Unsupported claims are visible in the editor UI. ``` --- ## 9.6 Parallel production node Purpose: Run production tasks in parallel after evidence matrix is ready. Parallel tasks: ```text Section scaffold generation Hero image prompt Diagram specifications Table specifications FAQ block SEO metadata Internal link suggestions Tone guide ``` Section scaffold output: ```json { "section_id": "uuid", "heading": "string", "draft_markdown": "string", "used_evidence_ids": ["uuid"], "unsupported_claims": ["string"], "suggested_visuals": ["string"] } ``` Acceptance criteria: ```text Every section from approved plan has one scaffold. Each section lists used evidence IDs. No scaffold silently introduces unsupported factual claims. High-risk unsupported claims block final approval until resolved. Parallel jobs can fail independently and be retried. ``` --- ## 9.7 Visual asset node Purpose: Create visual production instructions and optionally generate assets. v1 requirement: ```text Generate prompts and specifications. Store generated files if image tool is configured. Allow editor approval or replacement. ``` Asset types: ```text Hero image Inline diagram Table Flowchart Comparison matrix Architecture diagram ``` Diagram output example: ```json { "asset_type": "diagram", "format": "mermaid", "title": "AI content pipeline workflow", "diagram_code": "graph TD; A[Brief] --> B[Questions];", "alt_text": "Workflow showing article brief moving through planning, research, drafting, review, and publishing." } ``` Acceptance criteria: ```text Each visual has title, type, prompt or code, alt text, and status. Hero image prompt follows target site visual rules. Tables and diagrams are linked to article sections. ``` --- ## 9.8 Draft assembly node Purpose: Merge scaffolds, evidence, SEO metadata, and visual placeholders into one article draft. Markdown is the canonical editable draft format in v1. Rich preview renders Markdown/MDX for editor UX, but the stored approval artifact remains versioned Markdown plus metadata. Output: ```text Markdown draft Metadata Asset references Evidence references Unsupported claim warnings ``` Acceptance criteria: ```text Draft includes all planned sections. Draft includes title, meta title, meta description, body, FAQ block if applicable, and visual placeholders. Draft is saved as version 1. ``` --- ## 9.9 SEO review node Purpose: Apply target-site SEO rules. Checks: ```text Title length Meta title length Meta description length H1/H2 structure Keyword placement Internal links External citations Schema type Slug FAQ eligibility Image alt text Readability Duplicate headings ``` SEO output: ```json { "score": 82, "issues": [ { "severity": "medium", "field": "meta_description", "message": "Meta description is too long.", "suggested_fix": "Shorten to 150-155 characters." } ], "recommended_title": "string", "recommended_slug": "string", "schema_json": {} } ``` Acceptance criteria: ```text SEO report is visible in final review. Editor can accept or reject each SEO suggestion. Target-site SEO rules override global rules. ``` --- ## 9.10 Linguistic and tone review node Purpose: Check clarity, grammar, tone, brand voice, and consistency. Checks: ```text Grammar Spelling Sentence length Passive voice Jargon density Brand tone Forbidden phrases Repetition Weak claims Unsupported certainty CTA consistency ``` Acceptance criteria: ```text Issues are shown with severity and suggested rewrite. Editor can accept, reject, or edit suggestions. Final draft version stores accepted changes. ``` --- ## 9.11 Final approval gate Purpose: Stop before Git-backed publish commit creation. Final approval checklist: ```text Plan followed Evidence reviewed Unsupported claims resolved SEO metadata approved Images approved Tables and diagrams approved Internal links approved Frontmatter fields selected Content path selected Author selected Publishing mode selected ``` Acceptance criteria: ```text Publish commit cannot be created until final approval is complete. High-risk unsupported claims block final approval. Approval event stores actor, timestamp, article version, and selected publishing settings. ``` --- ## 9.12 Git-backed publish commit node Purpose: Create a content bundle and commit it directly to the configured production branch of a target Next.js site repository. Adapter interface: ```python class GitSitePublisher: def build_content_bundle(self, article: ArticleDraft, config: dict) -> ContentBundle: pass def run_transform(self, bundle: ContentBundle, site_workspace: str) -> TransformedBundle: pass def dry_run_preview(self, bundle: TransformedBundle) -> DryRunResult: pass def commit_to_production_branch(self, bundle: TransformedBundle) -> PublishCommit: pass ``` v1 publishing target: ```text Git-backed Next.js content repository Content bundle: Markdown/MDX file plus frontmatter JSON/YAML and referenced asset files Primary done state: PUBLISH_COMMIT_CREATED ``` Acceptance criteria: ```text System materializes versioned site publishing YAML/scripts into the checked-out site repository workspace. Admin-defined transformation scripts run directly on the runner host inside the checked-out site repository. Final editor approval is required before publishing. Best-effort dry-run validation uses the pipeline app's generic Markdown/MDX preview renderer, not the target site's actual build. Accepted v1 validation risk: site-specific build/runtime errors may reach production. System commits directly to the configured production branch. Target repository's existing CI/CD handles deployment after commit. Git conflicts or non-fast-forward pushes fail the publish step and require retry after refreshing from remote. If delayed deployment verification fails, system alerts Admin/Editor and leaves the production commit in place. System stores commit SHA, repository URL, branch, content bundle manifest, and publish status. ``` --- # 10. Multi-site configuration ## 10.1 Site config structure ```json { "site_id": "b2b_saas_blog", "name": "B2B SaaS Blog", "publishing_target": { "type": "git_next", "repository_url": "git@github.com:example/site.git", "production_branch": "main", "commit_author_name": "AI Content Pipeline", "commit_author_email": "pipeline@example.com" }, "editorial": { "audience": "B2B SaaS founders and content leads", "tone": "direct, expert, practical", "forbidden_phrases": ["revolutionary", "game-changing"], "preferred_structure": ["intro", "framework", "examples", "implementation", "conclusion"] }, "seo": { "title_max_chars": 60, "meta_description_max_chars": 155, "slug_style": "kebab-case", "schema_type": "Article" }, "sources": { "preferred_domains": ["official docs", "research papers", "government sources"], "excluded_domains": [], "minimum_sources": 5 }, "visuals": { "hero_ratio": "16:9", "style": "clean editorial illustration", "diagram_format": "mermaid" }, "publishing": { "content_format": "mdx", "content_path_template": "content/articles/{slug}.mdx", "asset_path_template": "public/articles/{slug}/{filename}", "frontmatter_mapping": { "title": "title", "description": "description", "date": "date", "author": "author", "category": "category", "tags": "tags" }, "transform_script_version_id": "uuid", "dry_run_renderer": "generic_mdx_preview" } } ``` ## 10.2 Acceptance criteria ```text Adding a new website must not require workflow code changes. Each site can define repository target, tone, SEO, visual, source, publishing, frontmatter, asset-path, and transformation rules. Workflow reads target site config at every generation step. ``` --- # 11. API endpoints ## Articles ```http POST /api/articles GET /api/articles GET /api/articles/{article_id} PATCH /api/articles/{article_id} DELETE /api/articles/{article_id} ``` ## Boundary questions ```http POST /api/articles/{article_id}/boundary-questions/generate GET /api/articles/{article_id}/boundary-questions PATCH /api/articles/{article_id}/boundary-questions/{question_id} POST /api/articles/{article_id}/boundary-questions/submit ``` ## Plans ```http POST /api/articles/{article_id}/plan/generate GET /api/articles/{article_id}/plans GET /api/articles/{article_id}/plans/{plan_id} POST /api/articles/{article_id}/plans/{plan_id}/approve POST /api/articles/{article_id}/plans/{plan_id}/request-revision PATCH /api/articles/{article_id}/plans/{plan_id} ``` ## Research ```http POST /api/articles/{article_id}/research/start GET /api/articles/{article_id}/research GET /api/articles/{article_id}/evidence PATCH /api/articles/{article_id}/evidence/{evidence_id} ``` ## Drafts ```http POST /api/articles/{article_id}/draft/assemble GET /api/articles/{article_id}/drafts GET /api/articles/{article_id}/drafts/{draft_id} PATCH /api/articles/{article_id}/drafts/{draft_id} ``` ## SEO and language ```http POST /api/articles/{article_id}/seo/review GET /api/articles/{article_id}/seo/report POST /api/articles/{article_id}/language/review GET /api/articles/{article_id}/language/report ``` ## Assets ```http POST /api/articles/{article_id}/assets/generate-specs GET /api/articles/{article_id}/assets PATCH /api/articles/{article_id}/assets/{asset_id} POST /api/articles/{article_id}/assets/{asset_id}/approve ``` ## Final approval ```http POST /api/articles/{article_id}/final-approval POST /api/articles/{article_id}/final-revision-request ``` ## Publishing ```http POST /api/articles/{article_id}/publishing/dry-run POST /api/articles/{article_id}/publishing/create-commit GET /api/articles/{article_id}/publishing/status GET /api/articles/{article_id}/publishing/commits ``` ## Site config ```http POST /api/sites GET /api/sites GET /api/sites/{site_id} PATCH /api/sites/{site_id} DELETE /api/sites/{site_id} GET /api/sites/{site_id}/publishing-config/versions POST /api/sites/{site_id}/publishing-config/versions POST /api/sites/{site_id}/publishing-config/versions/{version_id}/activate POST /api/sites/{site_id}/publishing-config/versions/{version_id}/rollback ``` ## Agent jobs ```http GET /api/agent-jobs GET /api/agent-jobs/{job_id} POST /api/agent-jobs/{job_id}/retry POST /api/agent-jobs/{job_id}/cancel ``` --- # 12. Prompt and instruction management ## Prompt modules ```text boundary_questions.md plan_generation.md research_brief.md evidence_matrix.md section_scaffold.md hero_image_prompt.md table_spec.md diagram_spec.md seo_review.md language_review.md draft_assembly.md publish_bundle.md ``` ## Prompt versioning requirements ```text Each prompt has version. Each agent job stores prompt version. Article history shows which prompt version generated each artifact. Admin can activate/deactivate prompt versions. ``` ## Prompt input format All prompts receive structured input: ```json { "article": {}, "target_site": {}, "boundary_answers": {}, "approved_plan": {}, "editor_notes": [], "evidence": [], "output_schema": {} } ``` --- # 13. Security requirements ## Credentials ```text 1. Do not store OpenAI or Anthropic API keys for generation in the app. 2. One managed Codex CLI subscription or enterprise access-token identity lives only on the runner host per environment. 3. Git credentials for target repositories are stored in a secrets manager or runner-local secure credential store. 4. The backend stores references to credentials, not raw secrets. 5. Agent generation jobs run in isolated workspaces with limited filesystem access. 6. Admin-defined publishing transform scripts are trusted host-level code and may run directly in checked-out site repositories. ``` ## Workspace isolation ```text 1. One workspace per agent job. 2. Workspace path must be generated by backend. 3. Runner cannot access sibling workspaces. 4. Job artifacts are copied to object storage after completion. 5. Workspace can be deleted after retention period. ``` ## Publishing safety ```text 1. v1 commits directly to the configured production branch. 2. Final approval is required before publish commit creation. 3. A generic Markdown/MDX content-shape dry run must pass before commit. 4. The dry run is best-effort and does not guarantee target-site build/runtime correctness. 5. Git conflicts and non-fast-forward pushes fail instead of auto-rebasing. 6. All publishing actions and script/config versions are logged. ``` --- # 14. Error handling ## Agent job errors ```text FAILED_SCHEMA_VALIDATION CLI_EXIT_CODE_FAILURE TIMEOUT MISSING_OUTPUT_FILE UNSUPPORTED_CLAIMS_FOUND SOURCE_RETRIEVAL_FAILED RESEARCH_ARTIFACT_UPLOAD_FAILED PUBLISH_DRY_RUN_FAILED GIT_CHECKOUT_FAILED GIT_COMMIT_FAILED GIT_PUSH_NON_FAST_FORWARD PUBLISH_VERIFICATION_FAILED ``` ## Retry rules ```text Plan generation: retry manually only Research: retry allowed Section scaffold: retry allowed per section SEO review: retry allowed Language review: retry allowed Publish commit creation: retry allowed only if no publish commit exists ``` ## Failure UI Each failure should show: ```text Job type Status Error category Error message Last successful step Retry button Logs for admin ``` --- # 15. Audit trail Every important action must be logged: ```text Article created Boundary questions generated Boundary answers submitted Plan generated Plan edited Plan approved Research started Evidence added Draft assembled SEO review completed Language review completed Asset approved Final approval granted Publishing dry run completed Publish commit created Delayed publish verification failed Site publishing config changed Site transform script changed Job failed Job retried ``` Audit event fields: ```json { "event_type": "PLAN_APPROVED", "actor": "user_id", "article_id": "uuid", "payload": {}, "created_at": "datetime" } ``` --- # 16. Development tasks ## Epic 1: Project foundation ### Task 1.1: Initialize repositories Deliverables: ```text Frontend app Backend app Runner service Shared schemas package Docker compose environment ``` Acceptance criteria: ```text Developer can run full local stack with one command. Frontend can call backend health endpoint. Backend can connect to Postgres. Runner service can receive test job. ``` ### Task 1.2: Database schema Deliverables: ```text Postgres migrations Core tables Indexes Seed site config ``` Tables: ```text users target_sites articles boundary_questions article_plans plan_sections evidence_items claims article_drafts assets workflow_events agent_jobs research_run_manifests publish_commits script_config_versions prompt_versions ``` Acceptance criteria: ```text Migrations run cleanly. Seed creates one sample target site. Article can be created and queried. ``` --- ## Epic 2: Article intake and boundary questions ### Task 2.1: Article creation API Acceptance criteria: ```text POST /api/articles creates article. Article status is ARTICLE_BRIEF_CREATED. Target site config is attached. WorkflowEvent is created. ``` ### Task 2.2: Boundary question generation Acceptance criteria: ```text System generates questions based on brief and site config. Questions are stored. User can edit answers. Required unanswered questions block plan generation. ``` ### Task 2.3: Boundary questions UI Acceptance criteria: ```text User can submit brief. User can answer generated questions. User can save partial answers. User can submit answers and continue. ``` --- ## Epic 3: Plan generation and approval ### Task 3.1: Codex runner MVP Acceptance criteria: ```text Backend can create AgentJob. Runner creates isolated workspace. Runner writes input files. Runner invokes Codex CLI. Runner captures logs. Runner validates output JSON. Runner updates job status. ``` ### Task 3.2: Plan generation workflow node Acceptance criteria: ```text Plan generated from brief, boundary answers, and site config. Plan saved as version 1. Workflow stops at PLAN_REVIEW_REQUIRED. ``` ### Task 3.3: Plan review UI Acceptance criteria: ```text Editor can review full plan. Editor can edit plan. Editor can request revision. Editor can approve plan. Approval creates WorkflowEvent. ``` --- ## Epic 4: Research and evidence ### Task 4.1: Research job Acceptance criteria: ```text Research runs only after approved plan. Research gathers sources for claims through explicit search/fetch scripts. Evidence items are stored. Each evidence item includes URL, title, source type, summary, quality score, retrieval timestamp. Research source sections and expanded metadata are uploaded to object storage. Research run manifest stores S3 keys, content hashes, source URLs, and artifact types. Insufficient acceptable evidence forces PLAN_REVISION_REQUIRED before drafting. ``` ### Task 4.2: Evidence matrix Acceptance criteria: ```text System maps claims to evidence. Unsupported claims are marked. Editor can view evidence by section. Editor can add or remove evidence manually. ``` ### Task 4.3: Evidence UI Acceptance criteria: ```text Editor can filter by section. Editor can filter unsupported claims. Editor can open source URLs. Editor can mark evidence as approved or rejected. ``` --- ## Epic 5: Parallel production ### Task 5.1: Section scaffold jobs Acceptance criteria: ```text One job is created per approved plan section. Jobs run independently. Each output includes draft markdown, evidence IDs, unsupported claims, and suggested visuals. Failed section can be retried alone. ``` ### Task 5.2: Asset spec jobs Acceptance criteria: ```text System creates hero image prompt. System creates table specs. System creates diagram specs. Each asset is linked to section or article. ``` ### Task 5.3: Draft assembly Acceptance criteria: ```text All successful section scaffolds are merged. Draft includes metadata and visual placeholders. Draft is saved as version 1. Workflow moves to SEO_AND_LANGUAGE_REVIEW_READY. ``` --- ## Epic 6: SEO and linguistic review ### Task 6.1: SEO review job Acceptance criteria: ```text SEO report generated from target site rules. Report includes score, issues, suggested fixes, metadata, slug, schema. Editor can accept or reject suggestions. ``` ### Task 6.2: Language review job Acceptance criteria: ```text Language report includes grammar, clarity, tone, repetition, and forbidden phrase issues. Suggestions are linked to exact draft locations. Editor can accept or reject suggestions. ``` ### Task 6.3: Final review UI Acceptance criteria: ```text Editor sees draft, assets, SEO report, language report, and evidence warnings. Editor can edit final draft. Editor can approve final version. Workflow stops until final approval. ``` --- ## Epic 7: Git-backed Next publishing ### Task 7.1: Git publisher interface Acceptance criteria: ```text GitSitePublisher interface implemented. Publisher builds Markdown/MDX content bundles with frontmatter and referenced assets. Publisher materializes versioned YAML config and transform scripts into the site repository workspace. Publisher captures transform logs and output manifest. ``` ### Task 7.2: Publish commit creation Acceptance criteria: ```text Publish commit creation requires final approval. Generic Markdown/MDX dry-run preview must pass. System commits directly to the configured production branch. Non-fast-forward push or conflict fails the publish step. Commit SHA, repository URL, branch, and content bundle manifest are stored. Publishing event is logged. ``` ### Task 7.3: Publishing UI Acceptance criteria: ```text Editor can edit frontmatter, content path, author, category, and tags. Editor can run generic Markdown/MDX preview. Editor can click Create publish commit. Publishing result shows PUBLISH_COMMIT_CREATED, commit SHA, and opportunistic deployment status if discoverable. ``` --- ## Epic 8: Multi-site support ### Task 8.1: Site config CRUD Acceptance criteria: ```text Admin can create, edit, and deactivate target sites. Site config includes editorial, SEO, source, visual, publishing, Git repository, frontmatter mapping, asset path, and transform script rules. Every Admin config/script change stores author, timestamp, diff, and rollback target. ``` ### Task 8.2: Site-aware workflow Acceptance criteria: ```text Every generation step receives target site config. Different sites produce different tone, SEO metadata, visual specs, frontmatter, asset paths, and content bundles. No hardcoded website logic exists outside adapters. ``` --- ## Epic 9: Observability and operations ### Task 9.1: Job logs Acceptance criteria: ```text Each agent job stores stdout, stderr, status, duration, and error category. Admin can inspect logs. Sensitive values are redacted. ``` ### Task 9.2: Workflow history Acceptance criteria: ```text Article detail page shows workflow timeline. Timeline includes user actions, system actions, and agent jobs. ``` ### Task 9.3: Retry and cancel Acceptance criteria: ```text Admin/editor can retry failed jobs where allowed. Admin/editor can cancel queued or running jobs. Cancelled jobs do not update article state. ``` --- # 17. MVP milestone plan ## Milestone 1: Foundation Goal: ```text Local stack, database, article creation, target site config. ``` Exit criteria: ```text User can create article brief for one target website. Article appears in dashboard. ``` ## Milestone 2: Plan approval loop Goal: ```text Boundary questions, Codex runner, plan generation, plan review. ``` Exit criteria: ```text Editor can approve or revise an article plan. Workflow stops correctly before research. ``` ## Milestone 3: Research and evidence Goal: ```text Research job, evidence matrix, claim-source mapping, object-storage research artifacts. ``` Exit criteria: ```text Approved plan produces evidence matrix. Research run manifest points to S3 source-section artifacts. Unsupported claims are visible. Insufficient evidence returns workflow to plan revision. ``` ## Milestone 4: Draft production Goal: ```text Parallel section scaffolding, asset specs, draft assembly. ``` Exit criteria: ```text System creates full markdown draft from approved plan and evidence. ``` ## Milestone 5: Review and Git-backed publish commit Goal: ```text SEO review, linguistic review, final approval, generic preview dry run, Git-backed publish commit. ``` Exit criteria: ```text Editor can approve final article and create a production-branch publish commit. ``` ## Milestone 6: Multi-site extension Goal: ```text Second target website with different config. ``` Exit criteria: ```text Same workflow works for two websites through site config and versioned publishing YAML/scripts. ``` --- # 18. Definition of done for v1 ```text 1. User can create article from short description. 2. System generates boundary questions. 3. User can answer questions. 4. System generates article plan. 5. Workflow stops for plan approval. 6. Editor can revise or approve plan. 7. Research starts only after approval. 8. Evidence matrix is created. 9. Section scaffolds run in parallel. 10. Visual specs are generated. 11. Draft is assembled. 12. SEO review is generated. 13. Linguistic review is generated. 14. Workflow stops for final approval. 15. Publish commit is created only after final approval and successful generic dry-run validation. 16. At least one target website works end to end. 17. A second target website can be added through config. 18. Agent generation runs through Codex CLI subscription auth, not direct API-key calls. 19. Every major action is logged. 20. Failed jobs can be inspected and retried. 21. High-risk unsupported claims block final approval. 22. Research source-section artifacts are stored in S3 with manifest records in Postgres. 23. Publishing validation is explicitly best-effort content-shape validation only. ``` --- # 19. Technical risks and mitigations | Risk | Impact | Mitigation | | ----------------------------------------------------------- | -----: | --------------------------------------------------------------------------- | | Codex CLI output is malformed | High | Use schema files, JSON-only instructions, validation, retry | | CLI auth expires | High | Runner health check, admin alert, re-auth flow | | Long jobs block backend | High | Use queue and runner service, not request thread | | Research creates weak evidence | High | Source scoring, unsupported claim flags, plan revision loop, editor review | | Expanded research corpus creates storage/legal burden | Medium | Store artifacts in S3 with manifests and hashes; document indefinite retention | | Multi-site rules become messy | Medium | Strict site config schema, versioned YAML/scripts, audit diffs | | Direct production-branch commit contains site-specific error | High | Final approval, generic dry run, explicit accepted validation risk, alerts | | Git push conflict or non-fast-forward | Medium | Fail publish step and require retry after refresh | | Admin transform script damages runner/site workspace | High | Treat Admins as trusted code operators, audit every version, rollback | | Agent invents unsupported claims | High | Claim extraction, evidence mapping, unsupported-claim gate | | Claude Code subscription use creates compliance issue | Medium | Keep Claude optional and internal-only | | Prompt changes break output | Medium | Prompt versioning and schema validation | --- # 20. First development ticket ## Ticket: Build v1 foundation for AI Content Pipeline ### Goal Create the initial backend, frontend, database, and runner-service foundation for the article workflow. ### Scope Implement: ```text 1. Article creation 2. Target site configuration with Git-backed publishing fields 3. Workflow state storage 4. Agent job table 5. Codex runner test job 6. Dashboard list 7. Article detail shell ``` ### Backend deliverables ```text FastAPI app Postgres migrations Article model TargetSite model WorkflowEvent model AgentJob model ScriptConfigVersion model POST /api/articles GET /api/articles GET /api/articles/{id} GET /api/sites POST /api/sites POST /api/agent-jobs/test-codex ``` ### Frontend deliverables ```text Dashboard page New article form Article detail page Status badge component Target site selector Publishing status display ``` ### Runner deliverables ```text Runner service process Workspace creation Input file writer Codex CLI command execution Log capture Output validation placeholder Job status update ``` ### Acceptance criteria ```text Given a configured target site When a user submits a short article brief Then the system creates an article with ARTICLE_BRIEF_CREATED status And the article appears on the dashboard And the article detail page shows workflow history And the target site can store repository, production branch, and initial publishing YAML config And an admin can run a test Codex job from backend And the runner stores stdout, stderr, exit code, and job status ``` --- # 21. Recommended implementation order ```text 1. Database schema 2. Backend article/site APIs 3. Frontend dashboard and article form 4. Workflow event logging 5. Agent job model 6. Runner service 7. Codex CLI test job 8. Boundary question generation 9. Plan generation 10. Plan review gate 11. Research and evidence 12. Parallel scaffolding 13. Draft assembly 14. SEO and language review 15. Git-backed publishing adapter 16. Second target website config ```