1. Why I kept explaining the same background
I was using ChatGPT for everyday tasks and kept running into the same frustration. It could help me work through an algorithm, but before it could help with a simple reply, I had to explain who I was writing to, what we had already discussed, and which decisions mattered.
That background often existed already, scattered across messages, documents, and notes. I still had to find it and assemble it into a prompt. Moving from one app to another meant doing that work again.
Could a tool notice what I was working on and bring the relevant background to me before I went looking for it?
I built Pacific to explore that question. The desktop prototype, Pacific.app, uses the active screen and stored context to suggest notes beside the current task.
Finding related text was only part of the problem. How should it decide which earlier decision still applies? When should it use a lasting preference, and when should it follow what is happening right now? Those questions shaped the ranking system.
2. What counts as useful context?
For a reply to a colleague, an old message can tell me what we agreed on. A separate note might remind me to keep the reply concise or avoid sharing an unfinished budget. Both can matter, but they serve different purposes.
I separated those into evidence and instructions, called metacontext in the prototype.
Earlier messages, meeting notes, shared documents, and past decisions that help explain the task.
Writing preferences, relationship notes, and boundaries on what should be included or shared.
A reply to Alex
As an example, imagine replying to a collaborator named Alex. The screen gives Pacific a name, a subject line, and the latest message. A vision-language model turns that visible information into a description the retriever can use.
The recommender then searches MemoryStore and available documents for Alex's collaborator notes, an earlier decision, or a relevant project brief. It ranks those candidates using both the current task and stored history.
The suggested package might include a meeting note alongside an instruction such as "Confirm the deadline; do not share the draft budget yet." If the contact is new, the interface can ask whether to save a collaborator note.
Related does not always mean useful. The user can inspect the suggested cards and choose what to bring into the conversation.
3. Keeping it beside the work
If finding context required opening another dashboard, I would still be interrupting the task I wanted help with. I put the suggestions in a floating desktop overlay, close to the active window.
I built Pacific Desktop with an Electron shell and a Python FastAPI backend. The screen observation, ranking, and card interface fit together like this:
A few interface decisions
- An overlay that leaves room to work: The frameless window stays on top and is click-through where empty, so the uncovered area remains usable.
- Two ways to bring a card into a chat:
- Drag Card Body: Transfers the note as
text/plaininto a chat input or email draft. - Drag File Handle: Uses an Electron file promise for an OS-level file drag into a file-upload area.
- Drag Card Body: Transfers the note as
- A small origami companion: Its expressions show states such as
watching,searching,found, andempty. The interface uses the ranking signal to decide whether to present a suggestion or show an empty state. I wanted the assistant to be able to shrug when it had little to offer.
How the ranking works
The prototype's input signals, scoring equations, and adaptive weights, with the limits of the approach.
Open Technical Details
The prototype's input signals, scoring equations, and adaptive weights, with the limits of the approach.
A. From screen context to suggested notes
These notes describe the prototype design. The equations make its ranking choices explicit; the coefficients are heuristic settings that still need evaluation across different tasks.
The pipeline observes the active task, ranks candidate files, snippets, and instructions, and packages selected context for the destination tool.
Input Signals
The scoring function uses five representations of the task and its background:
- Current Query ($q$): Formed dynamically by concatenating the explicit prompt ask, the active file path, recently accessed files, recent conversation turns, and the
screen_analysisandscreen_commentaryproduced by the Vision-Language Model (VLM). - Long-Term User Profile ($u$): Powered by
MemoryStore, an offline profile derived fromFileGramthat encodes the user's longitudinal behavioral patterns, domain expertise, vocabulary tendencies, and content preferences. - Live Workspace ($w$): At each tick, the VLM inspects the user's screen and populates a structured telemetry schema (active focused buffer, typing velocity, edit intensity). This telemetry is merged with the VLM's qualitative screen narrative into a single dense workspace vector.
- Short-Term History ($h$): A fast background LLM synthesizes recent dialogue turns into a concise single-sentence summary capturing the core subproblem, combined with the last 3 raw dialogue turns and 5 recently touched files.
- Active Task ($\text{task}$): A task-semantic embedding reflecting real-time generation: screen analysis, active files, recent messages, and recent write/edit session events.
The Context Package
Rather than dumping raw text into a prompt, the pipeline synthesizes a structured context payload partitioned into three functional tiers:
| Output Layer | Technical Description & Behavioral Function |
|---|---|
| 1. Evidence Context | Concrete files, snippets, threads, records, prior outputs, and entities to attach. Keeps retrieval task-aware and session-aware. |
| 2. Metacontext | High-level behavioral directives for the downstream model: stylistic writing rules, strict retrieval policies, citation requirements, confidentiality constraints, permitted tool calling limits, and escalation rules. |
| 3. Packaging | Target-specific serialization instructions tailoring the prompt bundle for the destination host (e.g. Chat Assistant, Coding Agent, Email Copilot, Analyst Workspace, or Autonomous Workflow). |
B. Combining relevance, history, and the current task
At each time step $t$, I score candidate items $d$ using query relevance, stored preferences, the workspace, recent history, and the current task. A redundancy penalty reduces the score of items that repeat context already selected.
$$\begin{aligned} \text{score}(d) &= \alpha \, s_{\text{base}}(d) + \beta \cos(u, d) + \gamma \cos(w, d) \\ &\quad + \eta \cos(h, d) + \zeta \cos(\text{task}, d) - \lambda \, \text{redundancy}(d) \end{aligned}$$
Where $u$ is the user's long-term profile, $w$ is the live workspace embedding, $h$ is short-term history, $\text{task}$ is the immediate task profile, and $\alpha, \beta, \gamma, \eta, \zeta, \lambda$ are dynamic weights recomputed on every scoring pass.
1. Query Relevance
Query relevance balances exact keyword matching with semantic intent:
$$s_{\text{base}}(d) = 0.6 \, \text{BM25}_{\text{norm}}(\text{query}, d) + 0.4 \cos(q, d)$$
Where $q = \text{embed}(\text{query})$, and the query representation is synthesized across all real-time sensory channels:
query = current_ask + active_file + recently_accessed_files + screen_analysis + screen_commentary + recent_conversation_turns
The VLM's screen_analysis and screen_commentary add a description of the visible task when the typed query gives little detail. Their usefulness depends on how accurately the screen is interpreted.
2. Long-Term User Profile
I combine a profile computed at startup with relevant stored memories retrieved for the current query:
- Static Profile ($u_{\text{static}}$): Pre-computed at startup from
FileGramoffline analysis, capturing long-term architectural preferences, programming languages, and design patterns. - Episodic Blend ($u_{\text{episodic}}$): At scoring time, the system retrieves memory chunks from
MemoryStore. Only chunks exceeding a strict threshold $\cos(q, \text{chunk}) \ge 0.25$ are retained. The top three passing chunks are averaged into $u_{\text{episodic}}$, and blended:
The $0.25$ cosine threshold filters weak matches before blending. It is a prototype setting, not a guarantee that the retained memories are relevant.$$u = \frac{u_{\text{static}} + u_{\text{episodic}}}{2}$$
3. Dynamic Weight Assignment
Some weights change with workspace divergence, activity, and the available task representation. Others use fixed values in the prototype:
| Weight | Dynamic Formulation | Operational Behavior |
|---|---|---|
| $\beta$ · History | β = exp(-3δ) (halved to β *= 0.5 if task vector exists) |
Decays exponentially as workspace divergence ($\delta$) increases. |
| $\gamma$ · Workspace | γ = (1 - exp(-3δ)) · workspace_activity (+0.1 in production) |
Increases the workspace contribution as activity rises. |
| $\alpha$ · Base Query | α = 0.4 (+0.1 during exploration) |
Boosted during open-ended research; balanced during coding. |
| $\eta$ · Short History | η = 0.15 |
Maintains steady conversational continuity across immediate turns. |
| $\zeta$ · Task Profile | ζ = 0.35 if task vector exists, else 0.0 |
Adds relevance to the current task when a task vector is available. |
| $\lambda$ · Redundancy | λ = 0.25 + 0.25 · budget_used |
Increases penalty from 0.25 to 0.50 as the context window fills. |
C. Adjusting when the task changes
A lasting preference may be useful while coding and irrelevant when replying to an email. I added adjustments for task changes and the amount of context already selected:
1. Profile Drift Detection
When the user's immediate work diverges from their historical habits, the engine detects profile drift by evaluating the cosine distance between the long-term user centroid $u_{\text{centroid}}$ and the current workspace fingerprint $w_{\text{fingerprint}}$ after z-score normalization:
$$\delta = 1 - \cos\left(u_{\text{centroid}}, \, w_{\text{fp\_norm}}\right)$$
As drift $\delta$ rises, the historical weight $\beta = \exp(-3\delta)$ decreases, while the workspace contribution increases according to its activity term. This is intended to reduce the influence of history when the current task diverges from it.
2. Context Window Pressure
As candidate chunks are sequentially packaged into the prompt payload, the remaining token budget contracts. To prevent prompt bloat and redundant restatements, the redundancy penalty $\lambda$ scales linearly with consumption:
$$\lambda = 0.25 + 0.25 \cdot \text{budget\_used}, \quad \text{budget\_used} \in [0, 1]$$
The redundancy penalty rises from $\lambda = 0.25$ at 0% budget use to $\lambda = 0.50$ at 100%. Repeated context therefore becomes less attractive as the package fills.
3. Multi-Variant Short Query Expansion
A short query such as "fix retry bug" leaves several interpretations open. The prototype expands it into three candidate resource descriptions ($q_1, q_2, q_3$), scores items against each, and averages the results:
$$\text{final\_score}(d) = \frac{1}{3} \sum_{i=1}^{3} \text{score}(d, q_i)$$
D. What I would try next
Screen observation has limits: covered windows hide information, switching windows adds delay, and VLM inference costs time and compute. These are directions I would like to explore beyond the current prototype:
- Direct tool context: Integrations through MCP could expose documents, diffs, or terminal output when the connected tools support them. I would compare that information with screen-derived context, including cases where one source misses what the other can see.
- Learning from selection: Accepted, edited, and dismissed cards could provide feedback for a learning-to-rank model. I would need to check whether those actions reliably indicate relevance before using them to replace the heuristic weights ($\alpha, \beta, \gamma, \eta, \zeta, \lambda$).
- Suggesting new memory: A system could propose recurring preferences from earlier sessions. The open question is how to distinguish a lasting preference from a one-time instruction, and when to ask the user before saving it.
5. What the prototype leaves open
Pacific turns the original question into something I can try: notes appear beside the task, and I can decide which ones to bring into a chat. The ranking combines current activity with stored history and reduces repeated context as the package fills.
The scoring latency and token reduction below describe the prototype's ranking and context-size results. They do not establish how much it improves an assistant's answers, or include every source of delay in the desktop experience.
I would like to evaluate that next. How often does a suggested card contain the decision a user actually needs? How quickly does the ranking recover after a task switch? When does the assistant need to ask instead of inferring context from the screen?
The idea that still interests me is simple: the background for a task often exists already. I want to make finding and reusing it less work, while leaving the person able to see and choose what the assistant receives.