Нотатки: захист

Робочі нотатки дослідження з посиланнями.

Defenses against prompt injection via LinkedIn descriptions and similar third-party text

Scope: who is exposed, how the attack surface is structured (acting vs summarizing tools), platform-level defenses, builder-level defenses and their admitted limits, user-level hygiene and detection signals, and frameworks. All notes in English, with URLs and dates. Compiled 2026-10-01.

Key question 1: Attack surface inventory - which tools read LinkedIn-like profile text, and which act vs only summarize

Takeaway

The exposed population is broad: LinkedIn's own recruiter-facing AI agent (Hiring Assistant) and third-party recruiting copilots ingest profile, resume, and message text as untrusted input, and the agentic ones can send messages and schedule screens on the recruiter's account. Browser agents (OpenAI ChatGPT Atlas agent, Claude for Chrome) add a second path: they act inside the logged-in web UI on behalf of either recruiters or candidates, so profile text they read becomes instructions for a process with the user's credentials.

Cited findings

Inferences

Gaps

Key question 2: Platform-level defenses - what LinkedIn does and its limits

Takeaway

LinkedIn's documented defenses are policy and product-design measures: a Responsible-AI-driven quality framework with human-in-the-loop approval for agent actions, recruiter review of agent recommendations, and general content/trust policies for the platform. No LinkedIn-specific prompt-injection mitigation (content filtering for hidden instructions, injection classifiers) is publicly documented in the sources found, and platform policy alone cannot stop injection because the content itself is user-generated and legitimate-looking.

Cited findings

Inferences

Gaps

Key question 3: Builder-level defenses and why they are partial

Takeaway

Vendors' documented defenses are: RL training to ignore injected instructions, injection-detection classifiers over untrusted content, red teaming, privileged-action gating (human approval, logged-out modes, watch modes), and developer patterns (tool-result isolation, labeling provenance, JSON encoding, least privilege, sandboxing, output screening). All major vendors and researchers state these reduce but do not eliminate risk: Anthropic reports a 1% attack success rate against an adaptive attacker on its browser agent and calls prompt injection "far from a solved problem"; OpenAI's CISO calls it "a frontier, unsolved security problem"; independent researchers argue guardrail percentages are a failing grade and the only reliable end-user defense is not combining private data, untrusted content, and external communication.

Cited findings

Inferences

Gaps

Key question 4: User-level defenses - what recruiters, job seekers, and individuals can do

Takeaway

User-level defense is hygiene, not immunity: treat all profile, resume, and message text as untrusted data that can contain instructions; give agents minimal privileges (read-only, no credential access when possible, human approval for sends); prefer well-scoped tasks over broad delegations; keep humans in the loop for any message or export; and watch for concrete signals of hidden text (white-on-white text, image-embedded instructions, obfuscated or split payloads). Vendors themselves place part of the burden on users: OpenAI recommends logged-out mode and warns that broad vague requests are riskier, and Anthropic's own example of an attack is hidden white text in an email.

Cited findings

Inferences

Gaps

Key question 5: Frameworks and standards - OWASP, NIST, MITRE, recruitment-specific

Takeaway

The applicable frameworks exist and agree on the same control set: OWASP's LLM Top 10 2025 puts prompt injection at LLM01 with seven mitigation families; NIST's adversarial ML taxonomy (AI 100-2) treats LLM prompt injection as part of the attack taxonomy; MITRE ATLAS has concrete technique IDs for direct and indirect injection. No recruitment-specific standard for prompt injection was found.

Cited findings

Inferences

Gaps

Honest state of the art (summary line for the report writer)

Every primary source agrees prompt injection is mitigated, not solved: Anthropic "far from a solved problem" with about 1% residual attack success under adaptive attack in browser use (2025-11-24), OpenAI CISO "frontier, unsolved security problem" (2025-10-22), OWASP "unclear if there are fool-proof methods of prevention" (LLM01:2025), and the independent researcher consensus that avoiding the trifecta (private data + untrusted content + external communication) is the only reliable user-level defense (2025-06-16). LinkedIn's public materials describe human-in-the-loop and Responsible-AI policy controls for its recruiter agent but no injection-specific filtering; no source found promises a silver bullet for LinkedIn-derived untrusted text, and none should be implied. Checkpoint 2026-10-01: research complete, 23 tool calls used, all five key questions addressed with primary sources. Status: complete.

← До змісту дослідження