0% found this document useful (0 votes)
4 views18 pages

Project Structure

The GlobeLens AI Build Plan outlines a comprehensive 18-week project for a three-person engineering team to develop a news intelligence platform that aggregates and synthesizes articles from multiple sources. The document details the project roadmap, team structure, technology stack, and phases of development, including foundational planning, backend and authentication, core features, and launch preparations. Key features of the platform include AI synthesis, reliability scoring, bias detection, and collaborative analysis tools.

Uploaded by

Mohamed Benmouh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views18 pages

Project Structure

The GlobeLens AI Build Plan outlines a comprehensive 18-week project for a three-person engineering team to develop a news intelligence platform that aggregates and synthesizes articles from multiple sources. The document details the project roadmap, team structure, technology stack, and phases of development, including foundational planning, backend and authentication, core features, and launch preparations. Key features of the platform include AI synthesis, reliability scoring, bias detection, and collaborative analysis tools.

Uploaded by

Mohamed Benmouh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

GlobeLens AI

Build Plan & Technical Specifications

A complete engineering and project management guide


for a team of three, from Day 1 to production launch.

Version 1.0
Estimated duration 18 weeks (part-time) / 9
weeks (full-time)
Team size 3 people
Audience Founding engineering team

Confidential — Internal use only


GlobeLens AI — Build Plan & Technical Specifications v1.0

Contents

1 Product Overview 3
1.1 The Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.2 The Solution: GlobeLens AI . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.3 Actors & Permissions Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3

2 Project Roadmap 4
2.1 Phase 0 — Foundations & Planning (Weeks 1–2) . . . . . . . . . . . . . . . . . . 4
2.2 Phase 1 — Core Backend & Auth (Weeks 3–6) . . . . . . . . . . . . . . . . . . . 5
2.3 Phase 2 — Article Page & Core Features (Weeks 7–10) . . . . . . . . . . . . . . . 5
2.4 Phase 3 — Advanced Features (Weeks 11–14) . . . . . . . . . . . . . . . . . . . . 6
2.5 Phase 4 — Admin, Polish & Launch (Weeks 15–18) . . . . . . . . . . . . . . . . . 6

3 Team Structure & Division of Work 7


3.1 Person A — Backend & AI Lead . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
3.2 Person B — Frontend Web Lead . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
3.3 Person C — Mobile & DevOps . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
3.4 Shared Team Rituals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8

4 Technology Stack 8
4.1 Frontend — Web . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
4.2 Frontend — Mobile . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
4.3 Backend . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
4.4 AI & Intelligence Layer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
4.5 Infrastructure & DevOps . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
4.6 Collaboration Tools . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8

5 Data Model 9
5.1 Core design principle . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
5.2 Key tables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
5.3 Naming conventions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11

6 The Scraping & AI Pipeline 11


6.1 Pipeline steps . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
6.2 Reliability score formula . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
6.3 Scraping legal considerations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12

7 API Design 12
7.1 Conventions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
7.2 Endpoint groups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12

8 Diagrams to Produce in Phase 0 12

9 Where to Start — Ordered Priorities 13


9.1 Week 1 — The three non-negotiables . . . . . . . . . . . . . . . . . . . . . . . . . 13
9.2 Week 2 — Schema consensus meeting . . . . . . . . . . . . . . . . . . . . . . . . 13
9.3 Week 3 — Prove the pipeline with two sources . . . . . . . . . . . . . . . . . . . 14
9.4 Scalability decisions to make early . . . . . . . . . . . . . . . . . . . . . . . . . . 14
9.5 Three risks that will kill the project if ignored . . . . . . . . . . . . . . . . . . . . 14

10 Specification Document Structure 14

1
GlobeLens AI — Build Plan & Technical Specifications v1.0

11 Appendix: Environment Checklist 15

2
GlobeLens AI — Build Plan & Technical Specifications v1.0

1 Product Overview

1.1 The Problem


Reading the news today means opening four or five tabs for the same event, re-reading the same
paragraphs across sources just to find one exclusive detail. This creates three distinct pain points:

• Redundancy fatigue. Readers waste time on identical content spread across outlets.

• Trust uncertainty. There is no quick way to gauge how many credible sources corroborate
a claim.

• Context collapse. Stories appear without their history, making it hard to understand how
a situation evolved.

1.2 The Solution: GlobeLens AI


GlobeLens AI is a news intelligence platform that automatically aggregates articles from multiple
sources, synthesizes them into a single authoritative article using AI, and surfaces exclusive details
from each outlet — without making users visit them individually.
Core feature set
AI Synthesis
Multi-source scraping (Al Jazeera, CNN, BBC, etc.) processed by Claude to
produce one coherent article. Each paragraph is traceable to its source(s).
Reliability Score
A 0–5 score calculated from the number and credibility weight of sources corrob-
orating a claim.
Dual-View Toggle between a classic editorial layout and an interactive world map (Dual-
View) where articles are pinned geographically.
Deep-Dive Timeline
A chronological thread of all related articles beneath each story, giving the full
history of an event from day one.
Fact-Checker
Authenticated users paste a link; the platform analyses it against its aggregated
database and returns a credibility assessment.
Bias Detection
An AI-powered indicator that labels each source’s political or editorial lean for a
given article.
Social Flux A sidebar showing trending social media reactions alongside traditional journal-
ism.
Collaborative Analysis
Authenticated users can post analysis threads beneath any article.

1.3 Actors & Permissions Matrix


2 Project Roadmap

3
GlobeLens AI — Build Plan & Technical Specifications v1.0

Feature Guest Auth. User Journalist Admin

Read home feed ✓ ✓ ✓ ✓


Geolocation-based feed ✓ ✓ ✓ ✓
Keyword search + filters ✓ ✓ ✓ ✓
Read article page ✓ ✓ ✓ ✓
View timeline ✓ ✓ ✓ ✓
View reliability score ✓ ✓ ✓ ✓
Write comments — ✓ ✓ ✓
Post analysis threads — ✓ ✓ ✓
Fact-checker tool — ✓ ✓ ✓
Interactive map view — ✓ ✓ ✓
Receive topic notifications — ✓ ✓ ✓
Submit journalist content — — ✓ ✓
Access admin dashboard — — — ✓
Manage users & comments — — — ✓
Add / promote articles — — — ✓

Table 1: Permissions matrix by actor type

The project is divided into five phases. Each phase has a clear deliverable that the entire team
reviews before the next phase begins.

2.1 Phase 0 — Foundations & Planning (Weeks 1–2)


Goal: No code is written until every team member has a shared, written understanding of the
data model, the API contract, and the full set of screens.

Phase 0 deliverables
• • OpenAPI specification (all endpoints, request/response schemas)
• • Finalised PostgreSQL schema with migration files
• • Scraping architecture decision (RSS vs. Playwright per source)
• • Wireframes for all screens in Figma (low-fidelity first)
• • Design system tokens (colours, spacing, typography)
• • [Link] monorepo initialised with ESLint, Prettier, TypeScript
• • GitHub repository with branching strategy (main/staging/feature)
• • Docker Compose for local development (API + DB + Redis + Elasticsearch)
• • GitHub Actions CI pipeline (lint + test + build)
• • React Native + Expo project initialised

Critical rule for Phase 0: The OpenAPI spec is Person A’s most important output
and is due by end of Week 1. Persons B and C cannot make final interface or mobile
decisions until they know the exact shape of every API response. Person B works from a
[Link] mock file until the real API is ready.

4
GlobeLens AI — Build Plan & Technical Specifications v1.0

2.2 Phase 1 — Core Backend & Auth (Weeks 3–6)


Goal: A working scraping pipeline that produces synthesised articles, stored in the database,
accessible via a secured REST API. The foundation everything else builds on.

Person A — Backend & AI Lead


• Authentication service: JWT (access + refresh tokens) + Google OAuth
• User registration, login, password reset endpoints
• Scraping pipeline: RSS ingestion for 2–3 seed sources, Playwright for dynamic pages
• Topic clustering algorithm (keyword + time-window grouping)
• AI synthesis endpoint using Claude API (structured JSON output with source attribu-
tions)
• Reliability score engine (weighted source count formula)
• Articles CRUD API (GET /articles, GET /articles/:id)
• Background job queue with BullMQ (scrape jobs, AI processing jobs)

Person B — Frontend Web Lead


• Authentication pages: login, register, forgot password
• Home feed (Standard view) consuming mock/fixture data
• Topic onboarding flow (shown on first visit)
• Search UI with keyword input and filter panel (date, location, topic)
• Shared component library: ArticleCard, SearchBar, TopicPill, NavBar

Person C — Mobile & DevOps

• Mobile: auth screens (login, register)


• Mobile: home feed screen
• Redis deployed and configured for API response caching + rate limiting
• Elasticsearch cluster configured and index mapping defined
• Staging environment deployed to Railway/Render

2.3 Phase 2 — Article Page & Core Features (Weeks 7–10)


Goal: The flagship article page is complete. By the end of this phase you have a testable MVP
worth showing to real users.

Person A — Backend & AI Lead


• Paragraph-to-source attribution API (returns JSONB with paragraph/source mappings)
• Timeline API: ordered list of raw articles for a given topic
• Comments API (create, read, delete)
• Bias detection pipeline integrated into AI synthesis call
• Fact-checker endpoint: accepts a URL, analyses credibility against the database
• Search fully integrated with Elasticsearch (filters: date, geo, topic, source)

Person B — Frontend Web Lead


• Article page: rendered synthesised body from JSONB data
• Hover-to-source UI: paragraph hover reveals source attribution tooltip

5
GlobeLens AI — Build Plan & Technical Specifications v1.0

• Timeline component: horizontal date-ordered thread of related articles


• Comments section (display + authenticated post)
• Reliability score badge component
• Bias indicator component (per-source labels)

Person C — Mobile & DevOps


• Mobile: full article page with timeline
• Mobile: comments
• Push notification integration (Firebase Cloud Messaging)
• Notification preferences screen (per-topic opt-in/out)
• Monitoring: Sentry installed on all three surfaces (web, mobile, backend)

MVP milestone. At end of Week 10, pause and show the product to 5–10 real users. Run
structured interviews: can they find a story, read the article, understand where information
comes from, and read the timeline? Their feedback will reprioritise Phase 3.

2.4 Phase 3 — Advanced Features (Weeks 11–14)


Goal: Dual-View map, fact-checker page, social flux sidebar, and the journalist actor.

Person A — Backend & AI Lead


• Geo-tagged articles API (lat/lng centroid per topic)
• Social flux integration (Twitter/X filtered stream API or Nitter RSS)
• Journalist content submission endpoints
• Rate limiting middleware (per IP and per authenticated user)
• Abuse protection: spam detection on comments via Claude classifier

Person B — Frontend Web Lead


• Interactive map view using Mapbox GL JS (country pins, region clusters)
• Dual-View toggle (Standard layout ↔ Map layout)
• Fact-checker page (URL input, analysis result display)
• Social flux sidebar component
• Bias indicator expanded UI with source-level detail

Person C — Mobile & DevOps

• Mobile: map view (react-native-maps or Mapbox React Native SDK)


• Mobile: fact-checker screen
• Admin dashboard: authentication + statistics screen
• Load testing (k6 or Artillery): target 500 concurrent users at <400 ms p95
• CDN configuration for static assets (Cloudflare)

2.5 Phase 4 — Admin, Polish & Launch (Weeks 15–18)


Goal: Production-ready. Admin dashboard complete, accessibility audited, app stores submit-
ted, monitoring in place.

6
GlobeLens AI — Build Plan & Technical Specifications v1.0

Person A — Backend & AI Lead


• Admin API: user management (block/delete), comment moderation, source manage-
ment
• Promoted articles logic (admin-flagged articles surface higher in feeds)
• Full test coverage: unit tests (Jest/Vitest), integration tests (Supertest)
• API documentation published (Swagger UI at /api/docs)

Person B — Frontend Web Lead


• Admin dashboard UI: all CRUD screens for users, comments, sources, articles
• Accessibility audit: WCAG 2.1 AA compliance, keyboard navigation, ARIA labels
• SEO: Open Graph tags, structured data (JSON-LD), [Link]
• Performance: Lighthouse score ≥ 90 on all four metrics

Person C — Mobile & DevOps

• App Store (iOS) and Google Play Store submission and review
• Production environment hardening: secrets in environment variables, no hardcoded keys
• Backup strategy: daily PostgreSQL snapshots, 30-day retention
• Grafana + Loki dashboards: scraping success rate, AI processing latency, API p95
• Disaster recovery runbook documented

3 Team Structure & Division of Work

3.1 Person A — Backend & AI Lead


Person A owns everything server-side: the database, the scraping pipeline, the AI processing
layer, and the REST API. This is the critical path of the project. Every feature that Persons B
and C build ultimately depends on the API being correct and stable.
Core skills required: [Link] (Fastify) or Python (FastAPI), PostgreSQL, Redis, Playwright,
Claude API, BullMQ, Elasticsearch, REST API design, Docker.
Key responsibility beyond coding: Maintaining the OpenAPI spec as a living document.
Every time an endpoint changes shape, Person A updates the spec and notifies the team in the
#deployments channel before deploying.

3.2 Person B — Frontend Web Lead


Person B owns the entire web experience from design system to production deployment. They
also own the Figma file — the single source of truth for what every screen looks like.
Core skills required: [Link] (App Router), TypeScript, Tailwind CSS, React Query, Zustand,
Mapbox GL JS, Figma, web accessibility (WCAG).
Key responsibility beyond coding: Writing fixture/mock data that mirrors the OpenAPI
response shapes exactly, so development never stalls waiting for the backend.

3.3 Person C — Mobile & DevOps


Person C mirrors Person B’s screens in React Native (consuming the same API), maintains all
infrastructure, and owns the admin dashboard in Phase 3.

7
GlobeLens AI — Build Plan & Technical Specifications v1.0

Core skills required: React Native + Expo, Expo Router, Firebase FCM, GitHub Actions,
Docker, Railway or AWS, Sentry, Grafana, Loki.
Key responsibility beyond coding: Keeping the staging environment stable and the CI/CD
pipeline green. No broken builds on main for more than two hours.

3.4 Shared Team Rituals

Ritual Format & cadence

Daily standup 15-minute async text update on Slack: done / doing / blocked
Weekly review 60 minutes synchronous: merge PRs, update Kanban, review API changes
PR review Every PR requires one review before merge. Rotate reviewers weekly.
Branching main (production) ← staging ← feature/xxx
Deployment Staging: automatic on merge to staging. Production: manual trigger only.
ADRs Every significant architectural decision gets a one-page Architecture Decision
Record in Notion.

Table 2: Team rituals and conventions

4 Technology Stack

4.1 Frontend — Web

Technology Rationale

[Link] 15 (App Router) SSR for SEO-critical news content; file-based routing; BFF API
routes
TypeScript Mandatory for a 3-person team. Catches API contract mismatches
at compile time.
Tailwind CSS Consistent utility-first styling; no CSS conflicts across teammates
React Query Server state management, background refetch, stale-while-
revalidate for live news
Zustand Lightweight client state (auth session, UI preferences, topic selec-
tions)
Mapbox GL JS Interactive country map with cluster markers for Dual-View feature

Table 3: Web frontend technology stack

4.2 Frontend — Mobile


4.3 Backend
4.4 AI & Intelligence Layer
4.5 Infrastructure & DevOps
4.6 Collaboration Tools
5 Data Model

8
GlobeLens AI — Build Plan & Technical Specifications v1.0

Technology Rationale

React Native + Expo Cross-platform iOS/Android; shares business logic with web
Expo Router File-based routing mirrors [Link] — same mental model for the whole
team
Firebase FCM Push notifications for topic alerts; works on both iOS and Android
react-native-maps Map view on mobile (wraps Apple Maps / Google Maps natively)

Table 4: Mobile technology stack

Technology Rationale

[Link] + Fastify High-throughput REST API; faster than Express; excellent TypeScript
support
PostgreSQL Primary relational database; JSONB support for storing synthesis output
Redis API response caching, rate limiting, session store, BullMQ broker
BullMQ Job queue for scraping, AI processing, and notification dispatch
Elasticsearch Full-text search with geo, date, and topic filters
Playwright Headless browser scraping for dynamic news pages (CNN, Al Jazeera)
RSS parser Lightweight ingestion for sources that publish RSS feeds

Table 5: Backend technology stack

Technology Rationale

Anthropic Claude API (Sonnet) Article synthesis, bias detection, fact-checking, comment
moderation
Structured outputs (JSON mode) Forces Claude to return machine-parseable JSONB with
paragraph/source maps
Prompt versioning in PostgreSQL All prompts stored with a version number, enabling A/B
quality testing
BullMQ AI processing queue Prevents synchronous Claude calls on user requests; all
processing is async
Redis caching for synthesis Synthesised articles cached with TTL; Claude is not
called on every page load

Table 6: AI and intelligence layer

5.1 Core design principle


The central architectural decision in GlobeLens is the separation between raw articles (what
was scraped) and synthesised articles (what users read). Every other table exists to serve this
distinction.

5.2 Key tables


users id (UUID PK), email, password_hash, role (enum: guest/user/journalist/admin),
topics_following (UUID[]), created_at, deleted_at (soft delete).

9
GlobeLens AI — Build Plan & Technical Specifications v1.0

Technology Rationale

Vercel [Link] frontend hosting; zero-config, global edge CDN, preview


deployments per PR
Railway (or Render) Backend API + worker processes; simpler than AWS for a 3-
person team
GitHub Actions CI: lint → test → build → deploy on every PR merge to staging
Docker + Docker Compose Local development parity; all services run identically on every
machine
Sentry Error tracking on frontend, mobile, and backend
Grafana + Loki Metrics and log dashboards; alerts on scraping failure rate and
API latency
Cloudflare CDN for static assets and DDoS protection in front of the back-
end

Table 7: Infrastructure and DevOps stack

Tool Use

Linear Kanban board, sprint planning, issue tracking


Figma All wireframes, high-fidelity designs, and the shared component li-
brary
Stoplight / Swagger UI OpenAPI spec editing and publishing
Notion Architecture Decision Records, onboarding docs, meeting notes
Slack / Discord #backend, #frontend, #mobile, #deployments channels
[Link] Database ERD design and sharing

Table 8: Collaboration and documentation tools

raw_articles id (UUID PK), source_id (FK), url (unique), url_hash (for deduplication),
title, body_text, published_at, scraped_at, processing_status (enum:
pending/clustered/synthesised/failed).
topics id (UUID PK), name, slug (unique), lat (decimal), lng (decimal), country_code
(ISO 3166-1 alpha-2).
articles id (UUID PK), topic_id (FK), headline, summary (1-sentence), body_jsonb
(JSONB: array of {text, source_ids[]}), reliability_score (numeric 0–
5), bias_data (JSONB), is_promoted (boolean), created_at, updated_at.
article_sources
Join table: article_id + raw_article_id.
sources id (UUID PK), name (e.g. “BBC”), base_url, rss_url, credibility_weight
(numeric 0–1, hand-tuned), scraping_method (enum: rss/playwright).
timelines topic_id (FK), raw_article_id (FK), position (integer), published_at.
Unique constraint on (topic_id, raw_article_id).
comments id (UUID PK), article_id (FK), user_id (FK), body (text), is_blocked

10
GlobeLens AI — Build Plan & Technical Specifications v1.0

(boolean), created_at, deleted_at.


notifications id (UUID PK), user_id (FK), topic_id (FK), article_id (FK), sent_at,
read_at.
prompt_versions
id, purpose (enum: synthesis/bias/factcheck/moderation), prompt_text, version
(integer), is_active (boolean), created_at.

5.3 Naming conventions


• All table and column names in snake_case.

• All primary keys are UUIDs (version 4), never auto-incrementing integers.

• Soft deletes use a deleted_at TIMESTAMPTZ column, never physical deletion.

• All timestamps are stored as TIMESTAMPTZ (UTC).

• Boolean columns use the prefix is_ (e.g. is_promoted, is_blocked).

6 The Scraping & AI Pipeline

This is the most technically complex and most business-critical part of GlobeLens. It must be
built and proven working before the frontend is wired to real data.

6.1 Pipeline steps


1. Ingestion (runs every 30–60 minutes via cron job): Fetch RSS feeds for sources that publish
them. For sources without RSS, use a Playwright headless browser. Compute a url_hash
for each item. Insert new items into raw_articles with status pending. Duplicate URLs
are silently ignored.

2. Clustering: Group pending raw articles by topic using a two-step approach: (a) extract
named entities and keywords using Claude or a lightweight NLP library; (b) group articles
that share ≥2 named entities and were published within a 48-hour window into the same topic
cluster. Create or update a topics record and update raw_articles.processing_status
to clustered.

3. AI Synthesis: For each updated cluster, send all clustered article bodies to Claude with
the active synthesis prompt. Claude returns structured JSON:

• headline: string

• summary: 1-sentence string

• body: array of {text: string, source_ids: string[]}

• bias: array of {source_id: string, lean: string, confidence: float}

• reliability_score: float 0–5

Store this as JSONB in articles.body_jsonb.

4. Timeline update: Insert all raw articles for this topic into the timelines table, ordered
by published_at. This gives users the full chronological history of a story.

11
GlobeLens AI — Build Plan & Technical Specifications v1.0

5. Indexing: Write the synthesised article to Elasticsearch. Fields indexed: headline, summary,
topic name, country code, geo-coordinates, published date. Invalidate the Redis cache key
for any cached response containing this topic.

6. Notification dispatch: Query users who follow this topic. Enqueue one notification job
per user. Notification worker calls Firebase Cloud Messaging. Record sent notifications in
the notifications table.

6.2 Reliability score formula


n
!
X
score = min wi , 5
i=1

where wi is the credibility_weight of source i (hand-tuned value between 0 and 1), and the
sum is capped at 5. Example weights: BBC = 0.9, Reuters = 0.95, tabloid = 0.25. This formula
is explainable to users and can be recalibrated without any database migration.

6.3 Scraping legal considerations

Important: Prioritise RSS feeds — they are explicitly published for consumption and
carry no scraping risk. For HTML scraping, check each site’s [Link] before adding it
as a source. Respect Crawl-delay directives. Consider reaching out to 2–3 major outlets
for formal data partnerships; this also provides a credibility story for the product.

7 API Design

7.1 Conventions
• All endpoints are prefixed /api/v1/.

• Authentication uses Authorization: Bearer <access_token> header.

• All list endpoints are paginated: ?page=1&limit=20 query parameters.

• All responses include a meta object: {timestamp, version, requestId}.

• Errors follow RFC 7807 Problem Details format.

• HTTP status codes are used semantically (200, 201, 400, 401, 403, 404, 422, 429, 500).

7.2 Endpoint groups


8 Diagrams to Produce in Phase 0

These eight diagrams must be completed in Phase 0 and reviewed by all three team members
before any code is written. They are the team’s shared contract.

1. System architecture diagram (Excalidraw or [Link]). Shows all services: web app,
mobile app, API server, scraper workers, AI processor, PostgreSQL, Redis, Elasticsearch,
Firebase, Sentry, Grafana. Include directional data-flow arrows.

2. Database ERD ([Link]). All tables, all columns with types, all foreign key rela-
tionships, all unique constraints. This is the most reviewed document of the project. A bad

12
GlobeLens AI — Build Plan & Technical Specifications v1.0

schema costs weeks to fix later.

3. User flow — General Person (Figma or Whimsical). Shows both guest and authenticated
paths. Entry → optional onboarding → home feed → article → branches: comment / fact-
check / map / follow topic / timeline.

4. User flow — Admin (Figma or Whimsical). Login → dashboard → user management /


article promotion / source management / comment moderation.

5. Scraping & AI pipeline flow (Excalidraw). Cron trigger → RSS/Playwright → dedu-


plication → topic clustering → Claude synthesis → reliability scoring → PostgreSQL →
Elasticsearch indexing → notification dispatch.

6. OpenAPI specification (Stoplight or YAML file in the repository). Every endpoint with
request body schema, response schema, auth requirements, and at least one example per
endpoint. This is Person A’s deliverable, due by end of Week 1.

7. React component tree (Figma component panel or a diagram). Defines the component hi-
erarchy: Layout → Page → Section → Card → Atom. Determines what lives in /components
(shared) versus /app/[page] (page-specific).

8. Wireframes for all screens (Figma). Low-fidelity first to validate flow, then high-fidelity
for hand-off. Screens: Home (Standard), Home (Map), Article, Search, Fact-checker, Topic
onboarding, Profile, Notifications, Admin dashboard (all sub-screens).

9 Where to Start — Ordered Priorities

9.1 Week 1 — The three non-negotiables

Person A: Publish the OpenAPI spec by end of Week 1. Even if it is incomplete or


placeholder, it must exist so Persons B and C can make decisions. Update it daily.

Person B: Produce low-fidelity wireframes for all screens by end of Week 1. Design
decisions made now cost nothing. The same decisions made during Week 10 cost days of
refactoring.

Person C: Set up the GitHub repository with the agreed branching strategy, ESLint,
Prettier, TypeScript strict mode, and a working Docker Compose that boots all services
locally. Every person must be able to run the full stack on their machine by end of Week 1.

9.2 Week 2 — Schema consensus meeting


All three people review the database ERD together. Pay particular attention to:

• The body_jsonb structure in articles — this is the most novel part of the schema and must
be agreed upon before synthesis is coded.

• The timelines table design — confirm the ordering and pagination approach.

• The prompt_versions table — confirm how active prompts are selected at runtime.

Write the first Alembic (Python) or Knex ([Link]) migration file together as a group exercise.

13
GlobeLens AI — Build Plan & Technical Specifications v1.0

9.3 Week 3 — Prove the pipeline with two sources


Do not build for scale in Week 3. Build for proof. Pick two sources (e.g. BBC RSS + Al Jazeera
RSS). Prove the full pipeline works end-to-end: scrape → cluster → Claude synthesis → stored
in PostgreSQL → returned by the API → rendered in the browser (even if ugly).
This end-to-end proof is the most important milestone of the project. Everything after it is
incremental.

9.4 Scalability decisions to make early


These are architectural choices that are cheap to make at the start and expensive to retrofit
later. Make them in Phase 0 or Phase 1, not Phase 3.

• Use a job queue from day one. Never call the scraper or Claude API synchronously inside
an HTTP request handler. Every slow operation goes through BullMQ.

• Cache synthesised articles in Redis. A synthesised article rarely changes. Cache it with
a TTL of 5–15 minutes. Invalidate on update. Claude is never called at request time.

• Paginate all list endpoints. No endpoint returns an unbounded collection. Default limit:
20. Maximum limit: 100.

• Store raw and synthesised content separately. Raw articles must never be deleted or
overwritten. You will want to re-run the AI pipeline with improved prompts against historical
data. This is only possible if the originals are untouched.

• Design the reliability score as a formula, not a stored procedure. The weights can
be updated in the sources table without any code deployment. The score is recalculated on
the next synthesis run.

• Use environment variables for all secrets. No API keys, database credentials, or Claude
API keys ever appear in source code or are committed to the repository. Use a .[Link]
file with placeholder values.

9.5 Three risks that will kill the project if ignored


1. Scraping legal risk. Some publishers prohibit scraping in their Terms of Service. Always
prefer RSS feeds. Check [Link] for every source before adding it. Respect rate limits
and crawl delays. Consider formal data partnerships.

2. Claude API cost runaway. A synthesis call on a 10-article cluster can cost non-trivial
amounts at scale. Set hard budget alerts in the Anthropic console from day 1. Cache
aggressively. Consider using a cheaper model (Haiku) for the first-pass clustering step and
Sonnet only for final synthesis.

3. Mobile being an afterthought. Every feature built in Phase 1 must be mirrored in React
Native during that same phase. Teams that defer mobile to the end always underestimate
the work and ship a degraded mobile experience or miss their deadline.

10 Specification Document Structure

Your internal specification document (hosted in Notion) should be structured as follows:

1. Executive Summary (1 page): what the product is and why it exists.

14
GlobeLens AI — Build Plan & Technical Specifications v1.0

2. Actors & permissions matrix (Table 1 from Section 1 of this document).

3. Feature specifications: one page per feature, each containing: description, user story,
acceptance criteria, out-of-scope notes, and open questions.

4. API contract: link to the OpenAPI spec on Stoplight / Swagger UI.

5. Data dictionary: every table and every column with type, constraints, and plain-English
description.

6. Non-functional requirements:

• Article page response time: < 400 ms at p95.

• Scraping cadence: every 30 minutes per source.

• Supported platforms: iOS 16+, Android 12+, evergreen browsers.

• Data retention: raw articles 90 days; synthesised articles indefinitely.

• Uptime target: 99.5% for web and API.

7. Architecture Decision Records (ADRs): one page per major decision. Format: Context
→ Decision → Rationale → Consequences.

8. Glossary: define all domain terms (raw article, synthesised article, topic cluster, reliability
score, etc.) so all three team members use the same vocabulary.

11 Appendix: Environment Checklist

Use this checklist at the start of the project to confirm every person’s local environment is
correctly configured.

□ [Link] 20+ installed (node –version)

□ Docker Desktop installed and running (docker compose up)

□ Repository cloned and .env file created from .[Link]

□ npm install (or yarn install) succeeds with no errors

□ npm run dev starts the [Link] frontend at localhost:3000

□ API server starts and GET /api/v1/health returns 200 OK

□ PostgreSQL reachable at localhost:5432

□ Redis reachable at localhost:6379

□ Elasticsearch reachable at localhost:9200

□ Expo Go installed on a physical device or simulator running

□ Anthropic API key added to .env and a test synthesis call returns successfully

□ GitHub Actions workflow passes on a test PR

□ Sentry DSNs configured (one per surface: web, mobile, backend)

15
GlobeLens AI — Build Plan & Technical Specifications v1.0

□ Each person has been added to: GitHub org, Linear workspace, Figma team, Notion workspace

GlobeLens AI — Build Plan v1.0 — Confidential

16
GlobeLens AI — Build Plan & Technical Specifications v1.0

Method Path Description

Authentication
POST /auth/register Create a new user account
POST /auth/login Return access + refresh to-
ken pair
POST /auth/refresh Exchange refresh token for
new access token
POST /auth/logout Invalidate refresh token

Articles
GET /articles Paginated feed (supports
geo, topic, date filters)
GET /articles/:id Single synthesised article
with body JSONB
GET /articles/:id/timeline Ordered list of raw articles
for the topic
GET /articles/:id/sources Source attribution data for
the article

Search
GET /search Full-text search
(?q=&from=&to=&country=&topic=)

Comments
GET /articles/:id/comments All comments for an article
(paginated)
POST /articles/:id/comments Post a comment (auth re-
quired)
DELETE /comments/:id Delete own comment or ad-
min delete

Fact-checker
POST /factcheck Submit a URL; returns cred-
ibility assessment (auth re-
quired)

Topics & notifications


GET /topics List all topics
POST /users/me/topics/:id/follow Follow a topic (auth re-
quired)
DELETE /users/me/topics/:id/follow Unfollow a topic (auth re-
quired)

Admin
GET /admin/stats Platform statistics
GET / PATCH / DELETE /admin/users/:id User management
PATCH /admin/comments/:id Block or unblock a comment
17
POST / DELETE /admin/sources Add or remove a scraping
source
PATCH /admin/articles/:id/promote Promote or demote an arti-

You might also like