InterviewHack.ai
Start free
Jobs / LILT (Production)

AI Benchmark Engineer | Native Language Specialist - French (France) - Remote

LILT (Production) · France (Remote)Remotesenior

In short

  • ▸Ingeniero de pruebas de IA que crea desafíos reales en francés para evaluar modelos multilingües.
  • ▸Día a día: construir entornos de tareas, scripts de verificación y analizar fallos en entornos de terminal con datos en francés.
  • ▸Lo destacado: el trabajo se centra en pruebas auténticas sin traducción al inglés, evaluando verdadera robustez multilingüe.

High English proficiency

Apply on company site ↗Share on WhatsApp

In ~1 minute you get: who interviews you, the likely questions answered from your CV, and your CV tailored to this job. Free, no card.

🎧Land the interview? Bring the copilot. Our free extension listens to the live interview and flashes 3-4-word anchors from your resume and prep — glance, connect, talk. Get the extension →

What they ask for

  • ✓5+ años de experiencia en ingeniería de software.
  • ✓Fluidez nativa o casi nativa en francés con dominio gramatical y de registro.
  • ✓Excelente dominio del inglés técnico.
  • ✓Experiencia sólida con CLI/terminal y scripting estándar.
  • ✓Conocimiento profundo de procesamiento multilingüe: codificación, locales, Unicode, texto bidireccional.
  • ✓Capacidad para crear pruebas deterministas y verificar resultados con precisión.

Don't tick every box? That's normal — your free dossier shows your gaps and how to cover them in the interview.

PythonShell scriptingData processingTerminal/CLI workflowsCoding agentsUnicode normalizationLocale-dependent conventionsText I/OToolchain interoperabilitySafe string operations

Who should you write to at LILT (Production)?

Your free dossier identifies the people who'd interview you — their background, what they value, and how to reach out so you stand out before applying.

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows. We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches. Note this is a remote, freelance opportunity What You’ll Deliver Task Engineering: Evaluating Coding Agents. Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially, these assets must remain in the target language to genuinely measure multilingual handling. Prompting & Translation: finding failure points where AI does not work, in your native language Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary). Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus). Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity. Qualifications Experience: 5+ years of industry experience in software engineering. Background: Proven track record at leading technology companies and/or graduation from top-tier engineering universities. Language: Native or near-native fluency, with a deep understanding of its grammar, register, and phrasing rules. High English proficiency. Technical Stack: Strong proficiency in Python, standard shell scripting, and data processing. Workflow: Extensive experience with Terminal/CLI-based development workflows and a working familiarity with coding agents. Domain Expertise: Deep technical understanding of multilingual text processing pitfalls, including: Encoding/decoding robustness and Unicode normalization. Locale-dependent conventions (collation, casing, non-Gregorian dates). Text I/O, toolchain interoperability, and safe string operations. (For specific languages) Bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts. Why Collaborate with Lilt? Your schedule, your rules. As an independent contractor, work when you want, as much or as little as you want. No fixed hours, no check-ins, no micromanaging. Get paid quickly and fairly. We respect your time and your expertise. Competitive rates, prompt payments, no chasing invoices. Work on projects that actually matter . Contribute to cutting-edge AI and language technology that is shaping how humans and machines communicate. Be part of something bigger. Join a global community of linguists, subject matter experts, and language professionals who are advancing human knowledge together. Grow without limits. As a Lilt contractor you get access to diverse, innovative projects that expand your portfolio and sharpen your skills across industries and domains. Have fun doing what you love. We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows.

Don't apply unprepared

We research who's interviewing you, tailor your CV and rehearse you live — first one free.

InterviewHack.ai

Prepare for the exact interview: who's interviewing you, a tailored CV, and a real coach.

Product

JobsFree ATS checkerInterview-English checkSalary checkLATAM salary reportFree coursesBlogTailored CVSpoken practiceIt's free

Remote jobs

ReactPythonFull-StackLATAMArgentinaMexicoSee all →

Prepare

Spoken practiceFrontendBackendAI EngineerBy companySell with your CV

Company

For employersAboutContactPrivacyTerms

© 2026 InterviewHack.ai · Your CV is yours. Never used to train anything. · A product of IA-PTY

Similar open roles

Project Manager, Gaming Localization Production

LILT (Production) · Remote

→

Project Manager, Gaming Localization Production

LILT (Production) · Remote

→

Medical Translators - German into Estonian - Remote

LILT (Production) · Germany (Remote)

→

Talent Manager

LILT (Production) · Remote

→