Back to all experience

Alignerr

Freelance / Remote (Indonesia)

Current Role

AI Data Labeling & Code Expert

Sep 2025 - PresentFreelanceRemote (Indonesia)

Worked on AI model improvement projects involving data labeling, prompt optimization, and automation of dataset workflows. Contributed to stronger model accuracy through high-quality data preparation and prompt testing.

Associated Projects

Case studies from this role.

These projects are linked from the project workbook's Associated With column.

View all projects
Agentic AI Trace EvaluationMay 2026 - May 2026

Agentic AI Trace Collection & Evaluation (OpenClaw Traces Licensing Project)

Curated and evaluated multi-step agentic AI execution traces used for large language model training, benchmarking, and autonomous agent evaluation.

+7
  1. Validated execution traces, metadata, artifacts, tool interactions, and outcome annotations.
  2. Rejected fabricated, incomplete, or low-quality trajectories to improve dataset integrity.
Speech Data EvaluationApr 2026 - Apr 2026

Air Traffic Control Audio Transcription & Evaluation (ATC Project)

Audio Transcription QA ContributorAI Data Labeling & Code Expert / Alignerr

Validated aviation speech datasets by producing FAA-compliant transcripts for air traffic control communications, supporting speech recognition and language model training.

  1. Generated accurate aviation transcripts using standardized phraseology and communication protocols.
  2. Improved transcription quality through speaker identification, callsign normalization, and terminology validation.
Agentic AI EvaluationDec 2025 - Jan 2026

Agentic AI Scenario Quality Assurance & Evaluation (Agent Error Analysis Project)

Designed and validated agentic AI evaluation scenarios by reviewing task specifications, execution traces, evaluator logic, and scoring rubrics to improve benchmark reliability for autonomous AI systems.

  1. Validated agent execution traces against expected behaviors using reproducible evaluation criteria.
  2. Identified evaluator inconsistencies, ambiguous scenarios, reward-hacking risks, and specification gaps.
AI Coding BenchmarkOct 2025 - Dec 2025

Web Development Task Generation & AI Evaluation Framework (Perfect WebGen 2 Project)

Frontend Benchmark DeveloperAI Data Labeling & Code Expert / Alignerr

Developed production-like frontend engineering benchmarks used to evaluate large language models on realistic web development tasks and software engineering workflows.

+9
  1. Created reproducible development environments, task specifications, and reference implementations.
  2. Automated functional and visual verification using Playwright, Puppeteer, and browser-based testing.

Role Detail

What this experience proves.

Responsibilities pulled from the professional experience workbook and tightened into recruiter-readable evidence.

01

Reviewed and labeled AI-generated code outputs against task requirements.

02

Evaluated code functionality, logic, completeness, and test-case behavior across programming tasks.

03

Maintained labeling accuracy and consistency through careful analysis and project guideline adherence.

04

Provided feedback on common error patterns in AI code generation.

05

Identified ambiguous instructions, edge cases, and mislabeled examples to improve datasets.

06

Helped refine prompt structures for better model understanding of programming tasks.

aldyth.ai Brain
Email

Ask anything about Aldyth

Use the portfolio brain to answer questions about experience, projects, tech stack, referrals, or what he can help build.

Ask about the profile

Search across work history, projects, skills, and opportunity links.

Ideas for writing

Hiring
Portfolio knowledge base