vega

Security and privacy engineering

What Vega's code actually does today, with the file that proves each line — including the parts that are not good yet.

version 4.1effective 2026-09-21

1How to read this page

Vega is a small company at pilot stage. We have no certifications, no external audit, and no security team. Publishing a page that implies otherwise would be the easiest thing on this list and the most damaging, so this page does something narrower instead: every claim below names the file in Vega's own source that implements it, and every gap we know about is listed in section 12 rather than left out.

If you are evaluating Vega for a company, sections 3, 5 and 12 are the ones to read first. If a claim here turns out to be wrong, tell us at help@tryvega.tech and we will correct this page and say what changed.

2Personal life is filtered out at ingestion

Every path that can write a capture runs a topical gate before anything becomes a work moment. Content the gate reads as health, personal finance, family and relationships, or personal interest is dropped and never becomes a session, a signal, a score or a graph node (lib/topicGate.ts). The gate's model call reads up to 1,800 characters of the capture, with credential shapes removed and after a pattern screen on Vega's side; that call is the filter, so what it reads is not yet filtered. It has run on Amazon Bedrock since 5 September 2026; section 5 says how.

The gate fails closed. If the classifier errors or times out, the content is dropped rather than kept — the code used to fail open on one path, a production incident traced stored personal-medical content to exactly that path, and the behaviour was changed (lib/topicGate.ts:14-27).

Building software about health, money or sport is work and passes the gate. Discussing your own health, money or family is personal and does not. The line is in the classifier prompt, and it is a classifier, so it is not perfect.

3There is no personal-information redaction

This section exists because it is the single most likely thing for a reader to assume from section 2, and the assumption is wrong.

Vega has three filters and none of them is a personal-information filter. The topical gate in section 2 decides whether a capture is about your own private life and drops the whole thing if it is; it does not look for personal data inside content it accepts. The secret redactor in section 6 matches credential shapes and replaces them; a person's name is not a credential shape. The anonymiser (lib/anonymize.ts) strips company names, people's names, codenames, geographies and amounts — but it runs only on the path that writes a public profile face or an anonymised training row, never on the ordinary stored record.

The consequence is set out in section 10 of the Privacy Policy: work content in Vega routinely contains other people's personal data, put there by the person who pasted it, and Vega does not detect it, cannot find it on request, and has no route to reach the person it belongs to. Building a detector is real work and it is item one in section 13.

4Raw prompt and response text is off by default

Vega's normal record of a session is a structured summary — a title, a short capsule, and derived signals. Keeping the verbatim prompt and reply is a separate, per-category, opt-in setting, and every category ships defaulted to summary-only, with the personal category defaulted off entirely (lib/types.ts:301-311).

There is exactly one place in the codebase that may authorise a raw write, and it denies by default. It denies when there is no seat, no category, an unknown category, a lookup failure, a category level that is not exactly raw, a missing encryption key, or a writing source that is not on a three-item allow list (lib/rawAuth.ts:66-118).

When a raw write is authorised, the text is encrypted with AES-256-GCM in the application before the row reaches the database, with a fresh initialisation vector per field and the authentication tag stored alongside, so tampering fails to decrypt rather than returning garbage (lib/rawCrypto.ts:20-24, 61-83). If the encryption key is missing or the wrong length, the server refuses to store the raw text at all rather than storing it in the clear (lib/rawCrypto.ts:35-48).

5Model calls run on a provider's service, not on Vega's machines

Every model call Vega makes obtains its client from one factory, modelClient() in src/adapters/model/provider.ts, with one exception: the topic gate's direct route, classifyViaDirect() in lib/topicGate.ts, still posts to Anthropic's API with its own fetch. That route is not the one production runs, because the gate is on Bedrock. For each of the fourteen call paths the factory reads one exact-string switch (providerFor(), the path's own variable and then MODEL_PROVIDER; ENGINE_PROVIDER covers classifier, rollup and threads together; only the string "bedrock" selects Bedrock) and hands back either an Anthropic API client (directClient(), configured by ANTHROPIC_API_KEY) or an Amazon Bedrock client (bedrockClient(), built on @anthropic-ai/bedrock-sdk). No static AWS key is stored in Vercel: bedrockCredentials() in lib/awsCredentials.ts exchanges the hosting provider's signed identity token for one-hour AWS credentials on the role vega-vercel-bedrock, whose permissions are bedrock:InvokeModel and bedrock:InvokeModelWithResponseStream on named model and inference-profile ARNs for Claude Haiku 4.5, Sonnet 4.6 and Sonnet 5 in us-west-2 and ap-northeast-1 (scripts/aws/vercel-oidc-bedrock-role.sh), and bedrockRegion() in the same file supplies the region, us-west-2 unless BEDROCK_REGION says otherwise. The role admits the us., global and jp inference profiles; production uses us., which may serve a request in Oregon, Virginia or Ohio. The account is an AWS account held personally by a founder and used as Vega's, and one legacy IAM access key on an older user in that account, unused by the application, remains pending deletion by the founders. The hourly memory job runs as an AWS function in us-west-2 under its own execution role (lambda/memory-writer/deploy.sh).

Which path is on which route is read at call time and reported by /api/health under inference, per path, with the model id, from the build that carries the model factory (inferenceMap() in lib/opsHealth.ts, PR #241, open and not yet deployed at the time of writing). As of 7 September 2026 production runs two paths on Amazon Bedrock, both since 5 September 2026 at 15:51 UTC, the time of the first rows in Vega's model ledger with provider bedrock: topic_gate, the personal-life filter (lib/topicGate.ts), and memory_writer, the context memory writer (src/adapters/model/memoryWriter.ts). The other twelve each have their own switch, none has moved, and each runs on Anthropic's own API where it runs at all: classifier, rollup and threads (lib/classifier.ts), derive (lib/deriveSessionContext.ts), ask (lib/askVega.ts), coach (lib/coachDeep.ts), anonymizer (lib/anonymize.ts), capsule (lib/embeddings.ts), interests (lib/interests.ts), semantic_leaks (lib/semanticLeaks.ts), retag (src/loops/retag.ts) and leak_fix (app/api/coach/leak-fix/route.ts, which takes its client from the factory like the rest). Six of the fourteen have no live caller in the product today: rollup, semantic_leaks, retag, leak_fix, ask (its screen was deleted) and coach (its route has no screen calling it). Every attempt on either route writes one row to the model ledger with the path, the provider and the model id (recordModelCall() in lib/modelLedger.ts).

When a screen asks Vega a question about your own work, this is what happens. No screen in the product does today; the route exists (app/api/ask-vega/route.ts) and what follows describes the code. Vega resolves your seat from your session, never from anything in the request, and retrieves only your own moments and threads. Where you have opted into keeping verbatim text, it decrypts the turns behind those moments through the one owner-scoped read path, which is parameterised by your seat id and has no parameter through which another seat could enter (lib/askVega.ts). Then it posts that material through the ask path, today Anthropic's API and, since 9 September 2026, the model claude-sonnet-4-6 (modelClient("ask", { seatId }) in lib/askVega.ts), and the citations attached to the answer are resolved on Vega's side against what was actually retrieved, so an answer cannot cite a moment that is not yours.

What travels: your moment titles, contexts, work types and signals; the titles of your open threads; and decrypted excerpts of your own prompts and replies, capped at 600 characters per transcript turn (MAX_EXCERPT_CHARS in lib/askVega.ts) and shorter for a cited moment. What comes back goes to you alone. The coaching path (lib/coachDeep.ts) has the same shape, and likewise no screen calls it today.

The direct fallback is off by default. For a path on Bedrock, directFallbackFor() in src/adapters/model/provider.ts returns a direct client only when the path's own fallback variable, or MODEL_FALLBACK, is exactly the string "direct" (fallbackEnabled()), and only for the four reasons that mean Bedrock cannot answer at all (FALLBACK_REASONS: no AWS credentials, a failed credential exchange, a role that does not admit the model, a model the region does not serve); never for a rate limit or a timeout. It is wired at two call sites, the capsule writer (lib/embeddings.ts) and the anonymiser (lib/anonymize.ts), switched off by default (the path's fallback variable unset); both of those paths have been on Bedrock since 8 September 2026, so the fallback can fire for them and Vega has not turned it on. It is not wired on the topic gate, the memory writer or the classifier. Since 9 September 2026 every model path is on Bedrock, so Vega's code sends every model request to Amazon only. A fallback that fired would be a second ledger row on the direct route, so it is visible rather than silent.

6Secrets are redacted before storage

Capture payloads are scanned for credential shapes and redacted before anything is written: provider API keys, GitHub and Slack tokens, AWS keys, bearer tokens, JWTs, PEM and OpenSSH private keys, Google API keys and OAuth secrets, and password-assignment lines (lib/safety.ts:11-45). Every model call passes its text through a second redaction on the way out, whichever route it takes (redactForEgress() in lib/redact.ts, called at each call site before the request body is built).

This is defence in depth, not a guarantee. It is pattern matching, it is tuned to prefer false positives over false negatives, and it cannot catch a secret that does not look like one. It is also only about secrets — see section 3. Do not paste credentials into a session you expect Vega to capture.

7The two pools, and what Vega may learn from

Everything Vega might learn from is either the content pool — the words you and the model wrote — or the structure pool, which is seat, date, model, tokens, work type, score, the nine signals, iteration count and fidelity, and nothing else. The engineering rules differ per pool and per act, and section 3 of the Privacy Policy is the full statement of them.

Nothing trains on the content pool, and the reason is in the source. A constant in lib/trainingStore.ts, TRAINING_STORE_ENABLED, is read before the training store's connection details on every training path. While it is off the store cannot be opened, every write returns without touching a database, and the health page goes red if the connection details are set regardless. It is not an environment variable and not a setting: it changes only in a reviewed code change, and a test in the credential-free CI gate sets both connection details and proves the store still does not open. Behind it, two older walls remain as defence in depth: the seat-type gate that excludes company accounts under any plan (lib/trainingStore.ts, trainingEligible), and the database-maintained consent flag that no application code can write (sql/migrations/099_consent_records.sql).

The code can also derive three content-free records — where a moment sat in its session, which observable terminal events followed it, and which decisions a person made about it. They are behind the same switch and are not written either.

8Cross-person numbers have a floor of three

No number that describes more than one person is shown unless at least three people contributed to it. The floor is a single constant used by every surface that computes one — team views, team intelligence, team decisions and efficiency (lib/vega/kfloor.ts and its callsites; the canonical constant is ORG_AGGREGATE_K_FLOOR in lib/orgs.ts and migration 040 hard-codes the same floor in SQL).

Below the floor the surface states the reason and renders nothing. It does not render a zero, and it does not render a smaller number.

In Aggregate mode, per-person rows on team surfaces carry recency, not volume, because a count beside a name reads as a ranking however it is sorted. Direct mode is the deliberate exception: a company can choose to show its managers per-person volume, cost and shipped counts, and every member is told so in those words on the join screen before they accept. There is no third state where a figure appears without having been disclosed.

9Deletion and retention

Deleting your account destroys your personal seat and the work hanging off it through the database's own cascade rules, behind a typed confirmation. It is deliberately designed so that one person closing their account cannot destroy their company's data: company seats are released rather than deleted before the login is removed (app/api/account/route.ts:9-31).

Four things are deleted on a clock today, by a maintenance sweep (lib/orphanSweep.ts:300-325, lib/captureRefusals.ts:321-327):

WhatDeleted after
Undo tombstones for deleted momentsAt least 24 hours, at the next maintenance sweep
Partially uploaded encrypted raw slices7 days
Gate refusal traces30 days
MCP activity records14 days

10Access, logging and analytics

Product surfaces are authenticated and scoped to the seat that owns the data. Owner-only material — cited moments, raw text, growth scores, coaching, captures — is never assembled into a manager, team, public, recruiting or benchmark payload.

Vega staff actions on customer data are written to an append-only log that is content-free by contract: it whitelists the fields each action may record, so a caller cannot put a prompt, a summary, an email address or a magic link into it, and email addresses are hashed before they reach it (lib/adminAudit.ts:1-12, 32-62). Opening the internal capture view is one of the logged actions.

Consent is its own record and it is append-only at the database, not a flag: a grant is a row, a withdrawal is another row, and the effective answer is derived from the history every time it is asked (sql/migrations/099_consent_records.sql). Update, delete and truncate all raise for every role including the service role, so a consent record cannot be quietly edited by anyone, including us.

Vega runs no third-party analytics, no advertising or marketing trackers, no session recording and no heat mapping. There is no such dependency in the application at all — the only cookies Vega sets are a login session and four display preferences (theme, density, accent, reduced motion).

11The inspectable boundary — what you can see about who sees you

The design goal is that at any moment a person can see who can see what about them. This is where that stands, split into what is built and what is not, because the difference is the whole value of the claim.

Built. The privacy settings screen carries a "who can see any of this" panel with three live rows drawn from real state: Vega staff, with the access described in section 10; your team, showing whether your company's prompt visibility is on or off at that moment; and the public internet, showing whether your profile is live and how many items are on it (components/rebuild/settings/PrivacySection.tsx:189-229). Every change to the company visibility setting is written to a company audit log that any member can open and that is protected against edits and deletes in the database. Your public profile lists every fact you approved. Your export produces your own record and names inside itself what it leaves out.

Not built. Vega does not log reads of your work by other members, so it cannot tell you which of your moments a named teammate has actually opened. The staff log records that a page was opened, not which rows were on it, so it cannot tell you whether a specific moment of yours was read. There is no per-item answer to "who can see this one" — only the three category rows above. And nothing exists for a third party who appears inside your content, because Vega cannot identify them (section 3).

What would close it. One surface and one log: a per-item visibility inspector, reachable from any moment, that lists every audience which can currently reach that item and why; backed by a read-side access log that records team and staff reads with the entry written before the content renders, on the same fail-closed pattern the internal capture view already uses. Until both exist, "you can always see who can see what" describes three category rows, and this section is written so that nobody reads it as more than that.

12What Vega does not have

This is the complete list as of the effective date, not a selection.

  • No certifications. No SOC 2, no ISO 27001, no penetration test report, no external security audit.
  • No personal-information redaction. Section 3. Names, addresses, contact details, identifiers and health or financial details inside work content are stored as written.
  • No zero-retention agreement with either model company. Since 9 September 2026 session text goes to Amazon Web Services on every model step, under standard terms, at capture and, when a screen calls the ask or coaching path, again at query time; Anthropic's own API receives nothing from Vega's code. Amazon's published position for Bedrock is no storage by default and no access for the model provider; it is Amazon's statement, which Vega cannot verify, not an agreement. There used to be a third company, OpenAI, receiving the sanitised capsule for a search vector; that path was deleted on 19 August 2026 and embeddings now run inside Vega's own servers. One company fewer is a real reduction and it is not the same thing as a retention agreement, which Vega has with nobody. See the subprocessor list.
  • No read-side access log and no per-item visibility inspector. Section 11.
  • No signed Data Processing Agreement on offer. A draft exists and is with counsel; it is not something we can currently execute.
  • No retention limit on content. Four short clocks exist (section 9); the content itself has none.
  • No customer-visible access log. The staff action log exists and no customer can read it.
  • No customer-selectable data residency. Storage and the web tier run in Tokyo. Model inference does not: since 9 September 2026 every model step runs in the United States on Bedrock (requests go to us-west-2, Oregon, under a profile that can serve them in Oregon, Virginia or Ohio), and the hourly memory job runs in Oregon. Files a person chooses to keep are stored in Oregon too, on S3, encrypted at rest with SSE-S3, in a bucket that blocks all public access, readable by the owner only through fifteen-minute presigned links, and deleted with the file or the account; this is storage, not inference — no model reads them. Section 17 of the Privacy Policy explains what that means for a European or Turkish user.
  • No per-account encryption keys and no key rotation path. One key, held by us.
  • No paid tier, and therefore no code that enforces the paid-content wall. The rule is stated in the Privacy Policy and is currently true only because no seat is paid.
  • No benchmark, and no control that records an opt-in to one. Section 8.
  • No delete path into the training store. A withdrawal is recorded and queued, not executed. Nothing is written there today, which is the only reason that is survivable.
  • No organisation-level privacy controls. Every privacy setting belongs to the individual. A company administrator cannot turn capture off for their own company, and cannot see or change a member's settings.
  • No self-hosted or on-premise option.

13What we are working on, in order

Stated so that this page can be checked against reality later, not as a commitment or a delivery date.

  • A personal-information detector on the ingest path, and a decision about what it does when it fires — refuse, redact, or store and mark.
  • Zero-retention terms with Amazon Web Services and with Anthropic, which is now the difference between what section 5 says and what a customer expects it to say.
  • A decision on the remaining twelve model paths. Each has its own switch and none has moved from Anthropic's API to Amazon Bedrock; whether any does, one at a time or at once, is a decision the founders have not taken, and each move would be a change to the subprocessor list. The direct fallback described in section 5 stays off unless a person turns it on for a path that carries it; there is no per-customer control over it, only one switch per path for the whole service, and a customer cannot have it disabled or enabled for their seat alone.
  • A delete path into the training store, before that store is ever configured.
  • A read-side access log and the per-item visibility inspector in section 11.
  • A retention clock for content, with the number shown in settings.
  • Row-level logging of staff reads, and database-level immutability on the staff log in production.
  • Per-account key derivation and a rotation path.

14Reporting a vulnerability

Send it to help@tryvega.tech. We will acknowledge within five business days. We do not currently run a bug bounty and cannot offer payment. We will not pursue anyone who reports a finding in good faith, does not access, modify or retain data belonging to another person beyond what is needed to demonstrate the issue, and gives us a reasonable chance to fix it before publishing.