Back to all experience

EBIT

Freelance / Remote (Indonesia)

Current Role

AI Trainer Developer

Feb 2026 - PresentFreelanceRemote (Indonesia)

Contribute to training and evaluating large language models by solving programming tasks, reviewing AI-generated code, developing test cases, and designing benchmark tasks for agentic software engineering evaluation.

Associated Projects

Case studies from this role.

These projects are linked from the project workbook's Associated With column.

View all projects
Agentic AI BenchmarkJun 2026 - Sep 2026

MINA - Agentic AI Task & Benchmark Development

AI Trainer DeveloperAI Trainer Developer / EBIT

Designed and implemented production-grade software engineering benchmarks used to evaluate autonomous coding agents across realistic development environments and automated testing pipelines.

+6
  1. Developed reproducible software engineering tasks spanning backend systems, APIs, databases, debugging, and infrastructure workflows.
  2. Built automated validation pipelines, reference implementations, and scoring systems for objective benchmark evaluation.
Agentic AI BenchmarkMar 2026 - Apr 2026

OpenClaw RL - Agentic AI Evaluation & Benchmarking

AI Trainer DeveloperAI Trainer Developer / EBIT

Designed and evaluated multi-agent AI workflows using OpenClaw to benchmark reasoning, planning, debugging, tool orchestration, and software engineering capabilities of large language models.

+10
  1. Created realistic multi-stage benchmark scenarios reflecting production software engineering workflows.
  2. Standardized evaluation methodologies through reproducible execution environments and structured scoring frameworks.

Role Detail

What this experience proves.

Responsibilities pulled from the professional experience workbook and tightened into recruiter-readable evidence.

01

Evaluated AI-generated code for correctness, efficiency, readability, and adherence to task requirements.

02

Solved algorithmic and software engineering problems across multiple programming languages.

03

Developed comprehensive test cases to validate functionality, edge cases, and code robustness.

04

Provided clear technical explanations and human-readable rationales for evaluation decisions.

05

Identified logical errors, performance issues, and opportunities to improve AI-generated solutions.

06

Collaborated on annotation and evaluation workflows to enhance LLM coding performance and reliability.

07

Designed reproducible benchmark tasks, validation pipelines, and scoring workflows for agentic AI software engineering evaluations.

aldyth.ai Brain
Email

Ask anything about Aldyth

Use the portfolio brain to answer questions about experience, projects, tech stack, referrals, or what he can help build.

Ask about the profile

Search across work history, projects, skills, and opportunity links.

Ideas for writing

Hiring
Portfolio knowledge base