Back to all experience

Outlier

Freelance / Remote (Indonesia)

Completed Role

AI Coding Evaluator & RLHF Contributor

Mar 2026 - May 2026FreelanceRemote (Indonesia)

Contributed to LLM training and evaluation projects focused on coding, reasoning, RLHF-style review, response ranking, prompt refinement, and agentic AI validation.

Associated Projects

Case studies from this role.

These projects are linked from the project workbook's Associated With column.

View all projects
AI Coding EvaluationMay 2026 - May 2026

AI-Assisted Software Engineering Evaluation (Human Preferences)

Evaluated AI-generated software engineering solutions through production-style development workflows, code reviews, and human preference evaluation to improve large language model coding capabilities.

+6
  1. Reviewed competing AI-generated implementations for correctness, software architecture, maintainability, testing strategy, and production readiness.
  2. Performed production-style pull request reviews covering code quality, interface design, error handling, naming conventions, and developer experience.
LLM Reasoning EvaluationMay 2026 - May 2026

Meter Pavilion - AI Evaluation & Reasoning Analysis

Evaluated large language model reasoning, instruction following, and multi-step problem-solving through structured evaluation workflows to improve model alignment and benchmark quality.

+1
  1. Analyzed complex reasoning chains, edge cases, and instruction-following behavior across diverse benchmark scenarios.
  2. Produced structured evaluation feedback supporting model alignment, reasoning quality, and benchmark development.
Agentic AI RL EvaluationApr 2026 - May 2026

OpenClaw RL - Agentic AI Task Generation & Evaluation Framework

AI Coding Evaluator & RLHF ContributorAI Coding Evaluator & RLHF Contributor / Outlier

Designed, executed, and evaluated complex multi-stage agentic AI tasks in OpenClaw to benchmark LLM planning, tool orchestration, reasoning, safety, and artifact generation for reinforcement learning datasets.

+17
  1. Developed realistic multi-stage evaluation tasks that benchmarked autonomous AI agents across planning, reasoning, tool coordination, and execution quality.
  2. Improved benchmark reliability by enforcing reproducible environments, standardized evaluation rubrics, and verifiable artifacts for fair cross-model comparisons.

Role Detail

What this experience proves.

Responsibilities pulled from the professional experience workbook and tightened into recruiter-readable evidence.

01

Evaluated and ranked AI-generated coding responses for correctness, logic, efficiency, and instruction following.

02

Reviewed outputs across Python, JavaScript, PHP, Java, and Golang.

03

Performed response validation, bug identification, and edge-case analysis to improve model reliability.

04

Contributed to Claude AI training and OpenClaw evaluation workflows.

05

Verified test cases and dataset consistency for model training quality.

06

Worked on agentic AI and function-calling evaluation tasks with emphasis on workflow accuracy.

aldyth.ai Brain
Email

Ask anything about Aldyth

Use the portfolio brain to answer questions about experience, projects, tech stack, referrals, or what he can help build.

Ask about the profile

Search across work history, projects, skills, and opportunity links.

Ideas for writing

Hiring
Portfolio knowledge base