How we built a rules-based clipboard intelligence engine for Linux — local, deterministic, sub-millisecond for typical payloads (]{3,}|ftp://[^\s\"'<>]{3,}", re.IGNORECASE) _RE_EMAIL = re.compile(r"\b[A-Za-z0-9._%+\-]+@[A-Za-z0-9.\-]+\.[A-Za-z]{2,}\b") _RE_IPV4 = re.compile( # validates 0-255 per octet r"\b(?:(?:25[0-5]|2[0-4]\d|[01]?\d\d?)\.){3}" r"(?:25[0-5]|2[0-4]\d|[01]?\d\d?)\b" ) _RE_IPV6 = re.compile(r"\b(?:[0-9a-fA-F]{1,4}:){2,7}[0-9a-fA-F]{1,4}\b") # Pragmatic range matching MD5 (32), SHA-1 (40), SHA-256 (64); strict lengths planned for v2.2 _RE_HEX_HASH = re.compile(r"\b[0-9a-fA-F]{32,64}\b") # Known credential prefixes — the first line of secret detection _RE_SECRET_PATTERNS = [ re.compile(r"ghp_[A-Za-z0-9]{36,}"), # GitHub PAT re.compile(r"AKIA[0-9A-Z]{16}"), # AWS Access Key re.compile(r"sk-[A-Za-z0-9]{32,}"), # OpenAI / Stripe re.compile(r"xox[baprs]-[A-Za-z0-9\-]{10,}"), # Slack tokens re.compile(r"eyJ[A-Za-z0-9\-_]{10,}\.[A-Za-z0-9\-_]{10,}"), # JWT re.compile(r"-----BEGIN (?:RSA |EC |OPENSSH )?PRIVATE KEY-----"), ] # (Note: _RE_PATH, _CODE_KEYWORDS, and _CODE_BRACKETS are omitted here for brevity) # ── Pool-based keyspace entropy estimate ────────────────────────────────────── def _entropy_bits(text: str) -> float: """Estimate theoretical keyspace entropy: H = len * log₂(pool_size).""" pool = 0 if any(c in string.ascii_lowercase for c in text): pool += 26 if any(c in string.ascii_uppercase for c in text): pool += 26 if any(c in string.digits for c in text): pool += 10 # Any non-alphanumeric character (symbols, punctuation, or non-ASCII) adds 32 if any(c not in string.ascii_letters + string.digits for c in text): pool += 32 return math.log2(pool or 26) * len(text) # ── Main classification function ────────────────────────────────────────────── def detect_tags(content: str) -> list[str]: if not content or not isinstance(content, str): return [] tags: list[str] = [] text = content.strip() if _RE_URL.search(text): tags.append("#link") if _RE_EMAIL.search(text): tags.append("#email") if _RE_IPV4.search(text) or _RE_IPV6.search(text): tags.append("#ip") if _RE_PATH.search(text) and "#link" not in tags: tags.append("#path") if _RE_HEX_HASH.search(text): tags.append("#hash") # Secret detection: known prefix first, entropy gate as fallback is_secret = any(p.search(text) for p in _RE_SECRET_PATTERNS) if not is_secret and len(text) >= 16 and " " not in text and "\n" not in text: if _entropy_bits(text) >= 90: is_secret = True if is_secret: tags.append("#secret") # JSON: only attempt parse if it looks like an object or array if text.lstrip().startswith(("{", "[")): try: json.loads(text) tags.append("#json") except (json.JSONDecodeError, ValueError): pass # Code: multi-line with language keywords or dense bracket patterns if "\n" in text and len(text) > 40: if _CODE_KEYWORDS.search(text) or _CODE_BRACKETS.search(text): tags.append("#code") return list(dict.fromkeys(tags)) # deduplicate, preserve order Enter fullscreen mode Exit fullscreen mode Secret Detection: Two Layers The #secret classifier uses a two-step approach: Layer 1 — Known prefixes. Patterns like ghp_, AKIA, eyJ (JWT header) are unambiguous. Match any of them and the tag is assigned immediately, no entropy calculation needed. Layer 2 — Entropy gate. For everything else: single-line, no spaces, at least 16 characters, and a pool-based keyspace entropy ≥ 90 bits. The 90-bit threshold was tuned empirically to catch random high-entropy tokens, API keys, and strong bearer credentials while skipping short identifiers, dictionary words, and common camelCase variable names. Both conditions guard against a common pitfall: pure Shannon entropy on raw text produces too many false positives on long sentences or dense code snippets. The entropy function here is H = len × log₂(pool_size) — a measure of the theoretical keyspace, not character frequency distribution. Note on non-ASCII characters: Because any(c not in string.ascii_letters + string.digits ...) catches non-ASCII characters (e.g. Arabic script, accented letters, emoji), unicode text will trigger the +32 symbols branch. This heuristic is tuned primarily for ASCII credentials; multilingual script entropy is handled gracefully without throwing exceptions. Known Limitations Honest engineering requires talking about edge cases: Long Paths tagged as #secret: A Unix path like /home/kareem/StudioProjects/DotGhostBoard/main.py currently receives both #path and #secret — it satisfies the entropy gate because it has no spaces and mixes uppercase, lowercase, digits, and /. We've logged this as a known limitation; v2.2 will add #path as an explicit exclusion condition for the entropy fallback. Hashes flagged as #secret: A 32-character hex string (like an MD5 digest) uses digits + lowercase (pool = 36), yielding 32 × log₂(36) ≈ 165 bits, which clears the ≥ 90-bit gate. Consequently, MD5 and SHA digests receive both #hash and #secret; a dedicated check to prioritize #hash alone is slated for v2.2. 2. What the Tag Enables: Contextual Smart Actions Tags are displayed as colored pill chips on each card. More importantly, they drive Contextual Smart Actions — buttons that appear dynamically in the card toolbar based on what was detected. Tag Action What it does #link 🔗 Open Link QDesktopServices.openUrl() — no copy-paste to browser #json { } Format JSON Parses, re-serializes with 2-space indent, writes back to clipboard #email ✉ Compose Constructs a mailto: URI, hands off to default mail client #ip 📡 Copy IP Extracts just the IP from surrounding log noise #secret 🛡 → Vault Opens Vault drawer, prefills secret dialog, purges plaintext from history Here's the JSON formatter — the action developers use most: # ui/widgets/item_card.py (simplified) import json def _on_format_json(self) -> None: try: parsed = json.loads(self._raw_text) formatted = json.dumps(parsed, indent=2, ensure_ascii=False) QApplication.clipboard().setText(formatted) self._flash_badge("✓ Formatted") except json.JSONDecodeError: self._flash_badge("✗ Invalid JSON") Enter fullscreen mode Exit fullscreen mode One button. Minified JSON in, readable JSON back in your clipboard. No editor, no extra window. Figure 1: Auto-detected tags (#link, #ip, #code) and dynamic contextual smart actions (🔗 Open Link, 📡 Copy IP) on clipboard cards. 3. Pipeline Architecture The classifier only stays fast because it runs inside a well-isolated pipeline: Figure 2: The 4-phase sequential classification and dispatch pipeline. [System Clipboard] ──► ClipboardWatcher (background polling / events) │ ▼ ClipboardPipeline (pure policy layer) ├── App filter (skip blacklisted sources) ├── Self-paste guard (skip our own copy events) ├── AutoTagger (sync, pure, no I/O) └── Duplicate & pin policy │ ▼ HistoryService ──► SQLite (PRAGMA secure_delete) │ ▼ HistoryController ──► CardsView (lazy render) Enter fullscreen mode Exit fullscreen mode AutoTagger has zero dependencies on Qt, SQLite, or any network code. That isolation means: The UI thread is never blocked by classification. Unit tests need no display server or mocked widgets. You can swap, extend, or replace tagging rules without touching controllers or storage. 4. Measured Latency We wanted hard performance guarantees rather than assumptions. Here are timeit measurements on typical clipboard payloads running on an Intel Core i5-8350U @ 1.70GHz (Python 3.14.7, Linux x86_64), 1,000 iterations each: Figure 3: Microsecond-level classification benchmark measurements across typical clipboard payloads. Payload min µs median µs max µs ────────────────────────────────────────────────────────── URL 8.2 8.8 10.7 GitHub PAT 6.5 6.8 10.1 IPv4 address 7.1 7.5 12.7 MD5 hash 7.1 7.2 7.6 JSON (small, ~80 chars) 14.1 14.4 16.1 Code snippet 20.3 20.4 21.4 Plain English text 14.8 15.2 15.3 JSON (1KB payload) 80.3 81.8 82.6 Worst-case dense (7.8KB) 1380.0 1410.6 1460.0 Enter fullscreen mode Exit fullscreen mode The "< 1ms" guarantee holds for typical clipboard payloads (< 5KB), with small items (URLs, hashes, JSON) processing in 8–82 µs. At the 8KB cap, worst-case dense tokens (7.8KB of unspaced text forcing character-by-character entropy evaluation) measure ~1.4ms. Payloads exceeding 8KB automatically bypass the entropy gate to ensure the UI loop never hitches on huge paste events. Coming Up in Part 2 The productivity layer is half the story. The other half is what happens to sensitive data after it's detected. In Part 2, we cover the security architecture: 🔐 Standalone .vault encrypted backups — AES-256-GCM packages with AAD header authentication, zero plaintext on disk, PBKDF2 key derivation. ⚡ CSPRNG Password & Token Generator — Live keyspace entropy meter calibrated to 131.1 bits. ⏳ Secret Expiry Tracking — Calendar-based expiry with ⛔ EXPIRED and ⚠️ 3d left badges. 📜 Encrypted Version History — Last 3 versions retained, one-click atomic revert. 🧹 Coordinated Plaintext Sweep — Purging matching unencrypted history records upon Vault deletion. 📦 Try DotGhostBoard 2.1 100% free and open-source — Apache-2.0 License. GitHub: kareem2099/DotGhostBoard OpenDesktop / KDE Store: DotGhostBoard on OpenDesktop git clone https://github.com/kareem2099/DotGhostBoard.git cd DotGhostBoard python3 -m venv venv && source venv/bin/activate pip install -r requirements.txt python3 main.py Enter fullscreen mode Exit fullscreen mode Keyboard Shortcuts Shortcut Action Ctrl+Alt+V Summon / hide Dashboard (migrates to active workspace) Ctrl+Alt+Space Floating Spotlight Quick Search Ctrl+Shift+V Open / close The Vault encrypted drawer Questions about the classification heuristics, the entropy threshold choice, or the action pipeline? Drop them in the comments. 👻
DotGhostBoard 2.1 'Leviathan' — Smart Clipboard Tagging & Contextual Actions on Linux (Part 1)
Full Article
Original Source
Read the full article at Dev →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.