Browse Technical Opportunities
Rigorous RLHF, adversarial red-teaming, lock-free concurrency, and specialized domain evaluation backed by 100% pre-funded ledger escrow.
Llama-3-70B Multi-Turn Jailbreak & System Prompt Evasion Red-Teaming
Construct complex multi-turn adversarial prompt trees attempting to breach safety alignment constraints in Llama-3-70B and Claude 3.5 Sonnet. Craft nested roleplay scenarios, base64 payload encodings, and cypher-shifted instructions to trigger policy violations, log vulnerability vectors, and calculate automated toxicity scores.
Rust Unsafe Lock-Free SPMW Ring Buffer & Loom Concurrency Audit
Evaluate and benchmark paired LLM outputs implementing a lock-free Single-Producer Multi-Worker (SPMW) queue in Rust utilizing std::sync::atomic ordering primitives (Acquire, Release, SeqCst). Audit unsafe pointer dereferences, identify memory model race conditions under Loom model checking, and write detailed performance comparison reports.
SEC Form 10-K DuPont Financial Ratio Extraction & RAG Hallucination Audit
Audit complex LLM RAG outputs extracting consolidated balance sheets, ASC 842 lease liabilities, and 3-step DuPont Analysis return on equity (ROE) ratios from 50 SEC Form 10-K filings. Verify mathematical derivations, detect subtle financial hallucinations in footnoted debt covenants, and construct ground-truth evaluation datasets.
CRISPR-Cas9 Off-Target CFD Scoring & Gene Regulation RLHF Preference Ranking
Rank and evaluate paired LLM candidate answers analyzing gRNA off-target cleavage probabilities using Cutting Frequency Determination (CFD) scores. Cross-reference variant frequencies from gnomAD v4 and Ensembl canonical transcript annotations to score biological accuracy and pathway impact.
Cross-Border Data Transfer & SaaS SLA Risk Assessment LLM Benchmark
Audit and rate LLM legal reasoning outputs synthesizing multi-jurisdiction master service agreements (MSAs) containing EU-US Data Privacy Framework clauses, liability limitation caps, and GDPR Standard Contractual Clauses (SCCs). Highlight hallucinated statutory citations, jurisdictional conflicts, and ambiguous indemnity terms.
LLVM IR SIMD Vectorization & Register Allocation RLHF Performance Ranking
Perform side-by-side preference evaluation on candidate LLVM intermediate representation (IR) code generated for loop unrolling and AVX-512 vectorization passes. Benchmark register spilling, vector alignment, and memory alignment constraints to evaluate code model capabilities.
Variational Quantum Eigensolver (VQE) Circuit Decomposition & Decoherence Audit
Inspect and evaluate LLM-generated Qiskit code constructing Variational Quantum Eigensolver (VQE) ansatz circuits for molecular ground-state energy estimation. Verify CNOT gate count efficiency, zero-noise extrapolation (ZNE) error mitigation, and statevector state fidelity.
Circom ZK-SNARK Nullifier Double-Spend & Circuit Constraint Audit
Audit AI-generated Circom circuits designed for zero-knowledge anonymous payment verification. Identify missing signal constraint checks, unconstrained intermediate variables, and potential double-spend nullifier reuse vectors.
Swahili-Sheng Code-Switching & Dialectical Register Evaluation
Evaluate LLM synthetic dialogs and technical translations between formal Swahili, urban Sheng code-switching, and English. Flag noun-class agreement errors (M-/WA-, KI-/VI-), context-inappropriate register drift, and mistranslated technical terminology.
Next.js 14 App Router Server Action & Hydration Race Condition Benchmark
Grade and debug AI model solutions attempting to fix complex hydration mismatches, optimistic UI updates, and stale-while-revalidate cache synchronization bugs in Next.js 14 server action pipelines.