myRemoteJobs
Quality Assurance

GenAI Security Evaluation Engineer

Turing

Verified listing
Contract
RemotePosted 5d agoUp to $150/hr

Overview

Turing is a research accelerator for frontier AI labs and an enterprise AI deployment partner based in San Francisco. This role evaluates how effectively static analysis tools detect vulnerabilities in GenAI applications by building and annotating realistic vulnerable codebases for agent and RAG systems.

What You’ll Do

  • Build runnable agent/RAG repositories with realistic Sensitive Information Disclosure or Excessive Agency vulnerabilities across tool calling, memory, and MCP
  • Create vulnerable, fixed, and hard-negative code variants with minimal security-relevant differences
  • Trace and annotate security-relevant assets, data/action paths, controls, root causes, severity, and residual risk
  • Define authorization contexts and write deterministic tests validating vulnerable, fixed, and negative behavior
  • Recommend security controls and participate in calibration and peer review

Requirements

  • 5+ years in application/product security or security-focused software engineering, including secure code review
  • Experience with source-to-sink analysis, taint analysis, SAST, CodeQL, or Semgrep
  • Hands-on experience building LLM agents or RAG systems using frameworks such as LangChain, LlamaIndex, OpenAI/Anthropic SDKs, or MCP
  • Strong authorization knowledge including actors, trust boundaries, tenants, OAuth, IAM, identity, permitted data/actions, purposes, recipients, and document-level access control
  • Production coding experience in Python and/or TypeScript with strong judgment in distinguishing genuine findings from non-findings
  • At least 4 hours per day availability, minimum 20 hours per week, with 4 hours overlap with Pacific Time

Who Should Apply

This role suits security engineers who have built production systems and understand both threat modeling and code-level vulnerability analysis—particularly those who have done secure code review or worked on SAST tooling. It will frustrate you if you're primarily a penetration tester unfamiliar with static analysis, or if you lack real experience building LLM systems; this role requires you to author vulnerable code that realistically passes basic inspection.

Salary Insight

The hourly rate of up to $150 reflects the specialized expertise required (5+ years in security engineering plus LLM/agent system experience); the "up to" phrasing and mention of location adjustment suggests the actual rate depends on your evaluation and may vary. Clarify your expected rate early in conversations.

Required Skills

Cyber SecurityLLMCode ReviewsSASTPython

Compensation

Up to $150/hr

Location
Remote
Engagement
Contract
Posted
5d ago

Opens Turing’s listing on MyRemoteJobs — we don’t collect applications ourselves.

About Turing

Remote work platform connecting skilled professionals with AI labs and companies for software engineering, AI training, data, and specialist contract work.

Remote policy: Remote worldwide, but location eligibility varies by individual project. Many AI roles require minimum hours per week and U.S. Pacific Time overlap. Contractors and freelancers only, not employees.

How hiring works at Turing →

Will your CV get past the filter?

Most applications are rejected by software before a person ever reads them. Scan your CV using our ATS Checker tool, see exactly what's blocking it, then fix it in the CV Builder and apply knowing it'll get through.

Scan my CV

Get roles like this by email

Tell us what you do and we'll email you when a similar remote job is posted. No account needed.

Free. We email you as soon as a role matches — applying early matters. Only genuine matches, and unsubscribe in one click.

See our jobs more often in your search

Add MyRemoteJobs as a Preferred Source on Google. Free, takes seconds, and you stay on this page.

Application Tip

Lead with a concrete example of a vulnerability you've found in production code or a secure code review you've led—what the vulnerability was, how you traced it, and what control would catch it—especially if it involved authorization or data flow. Then connect this to LLM-specific patterns: mention a RAG or agent system you've built or reviewed, and name specific frameworks you've used. Generic security experience won't distinguish you; specificity on both traditional secure code review and hands-on LLM work will.

Sourced from MyRemoteJobs · verified against the original posting