# Hearem - Full Product Context for AI-Powered Development (Landing Page & More)

## 🧠 Overview

**Hearem** is an iPhone and iPad listening reader that combines Apple on-device voices with multiple online AI voice providers. It turns PDFs, webpages, scans, typed text, images, and documents into audio for multilingual listening and supported export workflows.

**Core Value in One Sentence**:

Hearem turns PDFs, webpages, and scans into audio you can listen to across languages and export—with Apple on-device voices and multiple AI providers.

**Founder**: Eddie Go

**Public Title**: Founder of Hearem

**Founder Bio**: Eddie Go is the founder of Hearem, an iPhone and iPad app that turns PDFs, webpages, and scans into audio with Apple on-device voices and multiple AI providers.

**创始人简介**：Eddie Go 是 Hearem 创始人。Hearem 是一款 iPhone 和 iPad 听读工具，支持 Apple 本地语音和多家 AI 声音，可将 PDF、网页和扫描内容变成音频。

**Founder Portrait**: `website/assets/basic/founder/web/eddie-go.webp`

**Portrait Source**: `website/assets/basic/founder/source/eddie-go.png`

**Public Brand Introduction**: Hearem makes written content easier to listen to. Its iPhone and iPad app combines Apple on-device voices with multiple online AI providers to turn PDFs, webpages, and scans into audio for cross-language listening, background playback, and supported export. Less screen. More room.

**品牌简介**：Hearem 让文字内容更适合被听见。它通过 iPhone 和 iPad 听读工具，把 Apple 本地语音和多家在线 AI 声音带进同一套体验，让 PDF、网页和扫描内容可以跨语言收听、在后台继续播放，并在支持时导出音频和字幕。少盯屏幕，多留一点空间。

**Canonical product facts**: [`docs/marketing/PRODUCT_FACT_SHEET.md`](./marketing/PRODUCT_FACT_SHEET.md)

---

## 🎯 Goal of Hearem

Hearem's mission is to be a thoughtful **mobile-first TTS (Text-to-Speech) app**—making written content easier to hear and return to. It's built to help users **save time, reduce screen fatigue, and keep learning or working** while they commute, walk, cook, or rest.

---

## 🎯 Core Value Proposition

### Problems It Solves

- Reading fatigue (especially from screens)
- Time constraints in consuming text-based content
- Lack of support for reading content while multitasking
- Accessibility for visually impaired users

### Benefits to Users

- Save time by listening instead of reading
- Improve learning productivity
- Consume content during other activities (e.g., commuting, walking, resting)
- Get access to multilingual, expressive voice outputs
- Keep data handling local where possible, with online processing made clear when a feature or provider requires it

### What Makes Hearem Unique (USP)

- Powerful yet minimalistic mobile TTS app
- Multiple voice providers in one app, plus free Apple system voices
- Cross-language listening with provider- and voice-specific capability boundaries
- OCR + web reader + document import + clipboard listener + share extension in one place
- Native Apple Vision OCR — fast on-device recognition where supported
- Intelligent AI-powered formatting & voice optimization
- Voice cloning and custom voice creation
- Standard audio export and SRT subtitle export when subtitles are available
- Cross-language and multi-voice support with emotional tone control
- iCloud sync & backup for cross-device continuity
- Local-first design — privacy by design

---

## 🚀 Features

### 1. **Text-to-Speech (TTS)**

- **High-Quality Speech**: Converts text into natural-sounding speech using multiple AI providers.
- **Voice Providers in One Place**: Compare multiple voice providers in one app; availability can vary by product rollout and runtime configuration.
- **Native Apple Voices (Free & Unlimited)**: Use Apple's built-in voices without premium character cost. They provide a reliable on-device option for supported languages.
- **Multilingual Support**: Includes English, Chinese (Simplified and Traditional), Spanish, French, German, Italian, Japanese, Korean, Russian, Arabic, and more.
- **Speech Styles & Emotions**: Multiple tones and prosody options with emotion control for premium voices.
- **Custom Voice Creation**: Create personalized voices by selecting gender, age, and characteristics (Beta).
- **Voice Cloning**: Upload audio samples to clone existing voices with multi-language support.
- **Premium Voice Controls**: Fine-tune pitch, volume, and emotional tone for all premium voices.
- **Dialect Conversion**: Transform text into various dialects before generating speech.
- **Clipboard Auto-Read**: Detects copied content and prompts auto playback for ultra-fast usage.

### 2. **OCR Text Recognition**

- **Image-to-Text Conversion**: Extracts text from camera or photo library images using Apple's native Vision framework — fast and on-device where supported.
- **Multi-Image OCR**: Supports batch image recognition (premium feature).
- **Instant Playback**: Extracted text can be read immediately.

### 3. **Webpage Reading**

- **Smart Content Extraction**: Parses and reads main body text from URLs.
- **Image Text Extraction**: Option to extract and read embedded image text.
- **Reader Mode**: Strips out ads and navigation for smooth reading.

### 4. **Document Processing**

- **Document Upload & Processing**: Upload document files (PDF, TXT, MD, RTF, EPUB formats) and extract text content using AI technology.
- **Content Extraction**: Automatic extraction of title, description, and full text from documents.
- **Document History Management**: Track and manage all processed documents.
- **One-Click Conversion**: Apply extracted document content directly to text-to-speech.

### 5. **Offline & Background Playback**

- **Background Playback**: Voice continues even with screen locked or app in background.
- **Native Apple Voices**: On-device processing works offline for supported voices and languages.
- **Planned Offline Voice Packs**: Users will be able to download voices for offline use.

### 6. **Customizable Speech Settings**

- **Voice Selection**: Choose from various Azure voices (gender, accent, language) plus custom voices.
- **My Voices Library**: Organize and manage custom and cloned voices.
- **Speed and Pitch Controls**: Modify rate and tone for optimal listening.
- **Emotion Configuration**: Adjust emotional tone for premium and custom voices.

### 7. **Magic Editing Tools**

- **Summarize**: Condense long content for quick listening.
- **Main Idea Extraction**: Pulls out core messages.
- **Format Optimization**: Makes raw input more suitable for speech.
- **Translation**: Real-time translation with language selection.
- **Smart Split**: Automatically segments text into pages for parallel processing.
- **Dialect Conversion**: Convert text into regional dialects before speech generation.
- **Custom AI Editing**: Refine text with your own prompts.

### 8. **User Privacy**

- **Local-first data handling**: Apple voices and OCR can run on-device; online voices and other online features send the content needed to complete the request to the selected provider.
- **Manual Sync & Clear**: Cache and account data can be deleted by user.

### 9. **Customizable UI**

- **Minimalist & Clean**: Designed for speed and usability.
- **Multilingual Interface**: Supports 15+ languages for global reach.

### 10. **History, Playlist & Export**

- **Playback Records**: Save audio sessions locally with play progress tracking.
- **Resume Playback**: Pick up exactly where you left off in any audio.
- **Playlist Management**: Create and manage playlists for organized listening.
- **Album Management**: Group and organize audio records into albums.
- **Bookmarks**: Save favorite moments in audio tracks for quick access.
- **Real-time Subtitles**: Synchronized subtitle display during playback (Premium).
- **Sleep Timer**: Auto-stop playback after a set duration.
- **Segment Repeat**: Loop specific sections for study or review.
- **Export Audio/Text**: Share clips or content outside app with customizable filenames.
- **Subtitle Export**: Export subtitles when available.
- **Merged Audio Export**: Combine segmented audio into a single file for export.

### 11. **iCloud Sync & Backup**

- **iCloud Audio Sync**: Sync audio recordings across devices via iCloud.
- **Cloud Backup**: Back up created works to the cloud for recovery if lost locally.
- **Backup Management**: Manage backups with sync status indicators.

### 12. **Share Extension**

- **Cross-App Sharing**: Share content directly from Safari, Photos, Files, and other apps into Hearem using the system share sheet.

---

## 💰 Pricing Model

Hearem adopts a **freemium** model:

### Free Plan

- 10,000 characters (lifetime, not per month)
- 50 initial image scans
- 50 initial webpage reads
- Access to limited voice library
- Native Apple voices (free & unlimited)

### Standard Plan

- **Price**: $3.99/month or $35.99/year (US) | ¥20.00/month or ¥196.00/year (China)
- 1,000,000 characters/month
- 20,000 premium voice characters per monthly purchase cycle
- Unlimited OCR, web reading, and document processing
- Access to full voice catalog including custom voice creation
- Voice cloning capabilities
- Emotion control for all premium voices
- Smart Split up to 5,000 characters/page
- Advanced editing tools
- Multi-image OCR
- Ad-free experience

### Pro Plan (Planned — not yet surfaced in the app)

- Not displayed or sold in the current app
- Do not publish Pro pricing or benefits until the entitlement and customer-facing UI exist

### Flexible Add-ons

- **Voice Cloning**: Purchase independently of subscription
- **Premium Voice Characters**: Buy additional characters as needed

---

## 👤 Target Audience

### General Users

- Use case: Read books, documents, long articles hands-free
- Pain point: Reading fatigue, multitasking

### Students

- Use case: Listen to study notes, academic articles, summaries
- Benefit: Learn on the go, better retention

### Professionals

- Use case: Listen to emails, reports, market news while commuting
- Benefit: Save time, boost productivity

### Visually Impaired Users

- Use case: Access otherwise unreadable content
- Benefit: Improved accessibility, digital inclusion

### Content Creators

- Use case: Convert written blogs/news into audio
- Benefit: Repurpose content into audio formats

---

## 📱 Platform & Architecture

- **Platform**: iPhone and iPad app in public release; Android launch candidate work exists but is not public
- **Frontend**: React Native (via Expo)
- **Data**: Local SQLite via Drizzle ORM, iCloud sync optional
- **Backends**:
  - Azure Speech API
  - Minimax TTS
  - ElevenLabs
  - Fish Audio (OpenAudio S1)
  - Qwen Voice
  - Google Gemini Voices (registered, production disabled)
  - TikTok / Doubao TTS
  - OpenAI TTS (registered, App scope disabled)
  - Apple Intelligence (on-device)
  - Jina Reader API (for web parsing)

---

### 📱 Why Mobile? Mobile-Specific Benefits

1. **Ready when you are**
   As a mobile app, Hearem stays in your pocket—ready to read aloud books, notes, articles, or webpages the moment you need it, without needing a computer or extra hardware.

2. **Perfect for On-the-Go Usage**

   Hearem is optimized for commuting, walking, cooking, or relaxing. You can listen to longform content while multitasking—using time that would otherwise be idle.

3. **Lock-Screen & Background Playback**

   Voice continues even with the screen off, preserving battery life and reducing distractions.

4. **Fast Input via Mobile Actions**

   Copy from any app → Hearem detects clipboard → Instantly read aloud. This is only possible thanks to mobile OS-level clipboard APIs and app integration.

5. **Camera-Powered OCR**

   Scan physical books, handouts, or signage instantly using your phone's camera—extract and listen to text in seconds.

6. **Integrated Sharing & Workflow**

   Share content directly from Safari, WeChat, or email apps into Hearem using the mobile share sheet—streamlining your workflow without switching devices.

7. **Network-tolerant**
   Mobile-first design keeps local capabilities available where supported, while online features make their network dependency clear.

8. **Mobile OS Optimization**

   The interface is built around mobile interaction: touch-friendly controls, haptic feedback, one-handed navigation, voice control potential (Siri Shortcuts in future).

---

In short: **Hearem gives readers more choice over the voice they hear and more control over the audio they keep.**

---

## 🎨 UI / UX

- **Design**: Clean, distraction-free, monochrome palette with green highlights
- **Flow**:
  1. Open app
  2. Input (paste, scan, link)
  3. Tap play
  4. Adjust voice/speed
- **Special UX**:
  - Clipboard listener
  - Modal prompts for pasted URLs
  - Magic edit tools are modular and responsive

---

## 🔐 Privacy & Security

- **Local-first where supported**
- **Online features use the data needed to complete the request**
- **User-deletable data**
- **Explicit cloud opt-in for sync (planned)**

---

## ⚙️ Scalability & Performance

- Azure Speech backend supports high-concurrency use
- Smart local caching ensures smooth real-time playback
- Optimized for mobile network conditions

---

## 🔧 Upcoming Roadmap

- Community voice packs
- Offline voice pack downloads
- Android launch candidate and release readiness
- Advanced voice training algorithms
- Voice sharing capabilities
- Document format expansion (Word, PowerPoint, etc.)
- Collaborative audio creation features

---

## 🏷️ Brand Info

- **Name**: Hearem TTS: Text to speech
- **Tone**: Calm, smart, minimal, accessible
- **Logo**: Stylized ear
- **Color Theme**: Black / White / Green
- **Tagline**: "Less screen. More room." / "少盯屏幕，多留一点空间。"
- **Primary CTA**: "Start free" / "免费开始"
- **Secondary CTA**: "Preview voices" / "试听不同声音"

---

## 🖥 Landing Page Content Strategy

### Suggested Sections

1. Hero Section: logo + Title + subtitle + demo video/image + CTA
2. Problem & Solution
3. Core Features (cards with demo gifs/videos/vector animations)
4. Benefits
5. How It Works (3 steps)
6. Example voices and preview
7. Testimonials (with play button to listen to tts from the customer reviews)
8. Pricing table
9. FAQ
10. App Store button + QR code

### Sub-pages

1. Pricing
2. About
3. Privacy
4. Terms of service
5. Contact
6. FAQ

---

## ✅ Must-Have Info on Landing Page

- Free vs Standard comparison; do not surface Pro until the App implements it
- Download CTA (App Store badge)
- Data privacy commitment
- Voice samples or demo (if possible)
- Multi-provider AI voices + Apple free voices callout
- Smart Split + OCR + Web reader + Document import callout
- Language + voice support highlight
- iCloud sync & backup highlight

---

This is your master context file for landing page design, copywriting, onboarding tutorials, or investor decks. Use this to inform UI copy, section architecture, prompt engineering, and assistant queries inside Cursor AI IDE.
