Back to featured work

Complete Project Catalog

All projects, organized for fast evaluation.

NewestEvidenceFirst

A recruiter-friendly scan of AI, data, banking automation, and full-stack work. Each card keeps the story tight: context, stack, and measurable engineering contribution.

Agentic AI BenchmarkAgentic AI Trace EvaluationAI Coding EvaluationLLM Reasoning EvaluationAI Music ProductAgentic AI RL EvaluationSpeech Data EvaluationAgentic AI EvaluationAI Training DataAI Coding BenchmarkETL AutomationBanking AutomationFinancial Data AutomationTransaction Reporting AutomationCustomer Journey AnalyticsFintech OperationsMarketing Data PlatformAirport Security OperationsE-Commerce Operations
Agentic AI BenchmarkJun 2026 - Sep 2026

MINA - Agentic AI Task & Benchmark Development

AI Trainer DeveloperAI Trainer Developer / EBIT

Designed and implemented production-grade software engineering benchmarks used to evaluate autonomous coding agents across realistic development environments and automated testing pipelines.

  1. Developed reproducible software engineering tasks spanning backend systems, APIs, databases, debugging, and infrastructure workflows.
  2. Built automated validation pipelines, reference implementations, and scoring systems for objective benchmark evaluation.
  3. Improved benchmark reliability by refining task specifications, debugging execution environments, and strengthening evaluation criteria.
Agentic AI Trace EvaluationMay 2026 - May 2026

Agentic AI Trace Collection & Evaluation (OpenClaw Traces Licensing Project)

Curated and evaluated multi-step agentic AI execution traces used for large language model training, benchmarking, and autonomous agent evaluation.

  1. Validated execution traces, metadata, artifacts, tool interactions, and outcome annotations.
  2. Rejected fabricated, incomplete, or low-quality trajectories to improve dataset integrity.
  3. Strengthened AI training datasets by ensuring authentic planning, debugging, recovery, and tool-use behaviors.
AI Coding EvaluationMay 2026 - May 2026

AI-Assisted Software Engineering Evaluation (Human Preferences)

Evaluated AI-generated software engineering solutions through production-style development workflows, code reviews, and human preference evaluation to improve large language model coding capabilities.

  1. Reviewed competing AI-generated implementations for correctness, software architecture, maintainability, testing strategy, and production readiness.
  2. Performed production-style pull request reviews covering code quality, interface design, error handling, naming conventions, and developer experience.
  3. Generated structured human preference datasets used to improve LLM reasoning, coding quality, and instruction-following performance through RLHF workflows.
LLM Reasoning EvaluationMay 2026 - May 2026

Meter Pavilion - AI Evaluation & Reasoning Analysis

Evaluated large language model reasoning, instruction following, and multi-step problem-solving through structured evaluation workflows to improve model alignment and benchmark quality.

  1. Analyzed complex reasoning chains, edge cases, and instruction-following behavior across diverse benchmark scenarios.
  2. Produced structured evaluation feedback supporting model alignment, reasoning quality, and benchmark development.
  3. Improved evaluation consistency through standardized human-in-the-loop workflows for next-generation LLM systems.
AI Music ProductMay 2026 - May 2026
MoodTune AI - AI-Powered Music Recommendation Chatbot screenshotMoodTune AI

MoodTune AI - AI-Powered Music Recommendation Chatbot

AI Application DeveloperPersonal project

Built an AI-powered music recommendation platform that combines Gemini API and Spotify API to generate personalized playlists from natural-language conversations, moods, and genre preferences.

  1. Designed an end-to-end conversational AI application with real-time music recommendations.
  2. Integrated Gemini API with Spotify APIs and embedded playback for seamless user interaction.
  3. Built a modern Streamlit interface optimized for rapid iteration and AI feature experimentation.
Agentic AI RL EvaluationApr 2026 - May 2026

OpenClaw RL - Agentic AI Task Generation & Evaluation Framework

AI Coding Evaluator & RLHF ContributorAI Coding Evaluator & RLHF Contributor / Outlier

Designed, executed, and evaluated complex multi-stage agentic AI tasks in OpenClaw to benchmark LLM planning, tool orchestration, reasoning, safety, and artifact generation for reinforcement learning datasets.

  1. Developed realistic multi-stage evaluation tasks that benchmarked autonomous AI agents across planning, reasoning, tool coordination, and execution quality.
  2. Improved benchmark reliability by enforcing reproducible environments, standardized evaluation rubrics, and verifiable artifacts for fair cross-model comparisons.
  3. Enhanced AI training quality by identifying failures in reasoning, instruction following, safety, factuality, and tool usage for higher-quality reinforcement learning datasets.
  4. Contributed to robust evaluation pipelines for next-generation autonomous coding and agentic AI systems through structured task design and comprehensive performance assessment.
Speech Data EvaluationApr 2026 - Apr 2026

Air Traffic Control Audio Transcription & Evaluation (ATC Project)

Audio Transcription QA ContributorAI Data Labeling & Code Expert / Alignerr

Validated aviation speech datasets by producing FAA-compliant transcripts for air traffic control communications, supporting speech recognition and language model training.

  1. Generated accurate aviation transcripts using standardized phraseology and communication protocols.
  2. Improved transcription quality through speaker identification, callsign normalization, and terminology validation.
  3. Maintained dataset consistency across noisy audio, incomplete transmissions, and edge-case scenarios.
Agentic AI BenchmarkMar 2026 - Apr 2026

OpenClaw RL - Agentic AI Evaluation & Benchmarking

AI Trainer DeveloperAI Trainer Developer / EBIT

Designed and evaluated multi-agent AI workflows using OpenClaw to benchmark reasoning, planning, debugging, tool orchestration, and software engineering capabilities of large language models.

  1. Created realistic multi-stage benchmark scenarios reflecting production software engineering workflows.
  2. Standardized evaluation methodologies through reproducible execution environments and structured scoring frameworks.
  3. Identified planning failures, reasoning limitations, safety issues, and tool-use behaviors to improve future model performance.
  4. Produced benchmark datasets supporting reinforcement learning, autonomous agents, and LLM evaluation research.
Agentic AI EvaluationDec 2025 - Jan 2026

Agentic AI Scenario Quality Assurance & Evaluation (Agent Error Analysis Project)

Designed and validated agentic AI evaluation scenarios by reviewing task specifications, execution traces, evaluator logic, and scoring rubrics to improve benchmark reliability for autonomous AI systems.

  1. Validated agent execution traces against expected behaviors using reproducible evaluation criteria.
  2. Identified evaluator inconsistencies, ambiguous scenarios, reward-hacking risks, and specification gaps.
  3. Improved benchmark quality through stronger rubrics, trace validation, and standardized evaluation workflows.
AI Training DataNov 2025 - Nov 2025

Egocentric Video Data Collection for AI Training (Terminator Project)

AI Data Collection QA ContributorAI Data Labeling & Prompt Engineer / G2i Inc.

Produced and validated high-quality egocentric video datasets for multimodal AI training, ensuring compliance with strict recording, privacy, and annotation standards.

  1. Collected first-person activity datasets supporting computer vision and multimodal AI training.
  2. Performed quality validation against recording standards, privacy requirements, and reviewer guidelines.
  3. Improved dataset consistency through iterative quality reviews and submission refinements.
AI Coding BenchmarkOct 2025 - Dec 2025

Web Development Task Generation & AI Evaluation Framework (Perfect WebGen 2 Project)

Frontend Benchmark DeveloperAI Data Labeling & Code Expert / Alignerr

Developed production-like frontend engineering benchmarks used to evaluate large language models on realistic web development tasks and software engineering workflows.

  1. Created reproducible development environments, task specifications, and reference implementations.
  2. Automated functional and visual verification using Playwright, Puppeteer, and browser-based testing.
  3. Improved benchmark reliability by eliminating ambiguous prompts and strengthening evaluation criteria.
ETL AutomationMay 2025 - May 2025

Data Automation Report Cust Perisai Plus

Built an automated ETL pipeline that centralized operational reporting data for Cust Perisai Plus, improving reporting efficiency and data consistency.

  1. Designed and automated end-to-end ETL workflows for operational reporting.
  2. Reduced repetitive manual processing through scheduled data automation.
  3. Improved reporting consistency and accelerated access to business insights.
Banking AutomationApr 2025 - May 2025
Auto Report SV MCC (Sales Volume Merchant Category Code) screenshotSV MCC Automation

Auto Report SV MCC (Sales Volume Merchant Category Code)

Developed an automated reporting service that generates Merchant Category Code (MCC) sales analytics from large-scale credit card transaction data.

  1. Built scheduled Go-based ETL jobs to automate transaction aggregation by Merchant Category Code.
  2. Reduced manual report generation while improving reporting consistency and data accuracy.
  3. Delivered timely sales analytics to support business monitoring and portfolio management.
Financial Data AutomationApr 2025 - Apr 2025

Data Automation ENR

Developed an automated ETL service for Estimated Net Revenue (ENR) reporting, transforming financial data into centralized analytics datasets.

  1. Designed automated pipelines for collecting, transforming, and loading operational financial data.
  2. Improved report availability while eliminating repetitive manual processing.
  3. Implemented scheduled Go services to increase reporting reliability and consistency.
Transaction Reporting AutomationApr 2025 - Apr 2025

Data Automation JCB Report

Developed a transaction reporting automation service that centralized JCB credit card transaction data into enterprise reporting systems.

  1. Built automated ETL workflows for JCB transaction reporting.
  2. Reduced manual reporting effort while improving reporting accuracy and consistency.
  3. Delivered standardized transaction datasets for downstream analytics and business reporting.
Customer Journey AnalyticsOct 2024 - Mar 2025
Cust Activation Campaign & Customer Journey Monitoring screenshotCampaign Analytics

Cust Activation Campaign & Customer Journey Monitoring

Built a customer journey analytics platform that monitors the complete lifecycle of credit cards from issuance through activation and transaction activity.

  1. Developed end-to-end customer journey tracking from card issuance to first and last transactions.
  2. Automated backend data processing using Python to synchronize customer activity across banking systems.
  3. Integrated multiple internal services through REST APIs to support personalized activation campaigns and business analytics.
Fintech OperationsDec 2023 - Aug 2025
Billing App Fintech screenshotBilling App Home

Billing App Fintech

Developed a fintech billing and reconciliation platform that tracks partner billing activities, validates payment reports against transaction data, and improves financial reporting accuracy.

  1. Built billing workflows covering partner usage, reconciliation, payment validation, and reporting.
  2. Automated reconciliation by comparing external billing reports with internal transaction records.
  3. Reduced manual verification effort while improving billing accuracy and operational efficiency.
Marketing Data PlatformJan 2023 - Apr 2025
Database Marketing Analytics screenshotMarketing Analytics Home

Database Marketing Analytics

Developed and maintained an enterprise analytics platform that consolidated customer, transaction, and financial data into interactive dashboards for marketing and portfolio management.

  1. Built KPI dashboards covering customer acquisition, transaction volume, ENR, installments, and merchant analytics.
  2. Integrated multiple enterprise data sources into a centralized analytics platform.
  3. Enabled business users to make faster data-driven marketing and financial decisions through interactive reporting.
Airport Security OperationsFeb 2021 - Apr 2021

e-Bappi

Developed a web-based operational management system that digitized airport security workflows for prohibited, entrusted, and lost property management.

  1. Built a full-stack application replacing manual paper-based operational processes.
  2. Implemented secure item registration, tracking, and reporting with real-time status updates.
  3. Improved operational efficiency by centralizing airport security records into a searchable system.
E-Commerce OperationsSep 2020 - Dec 2020

e-Commerce Admin Dashboard

Fullstack Web Developer (Intern)Full-stack Developer / PT Compro Kotak Inovasi

Developed a responsive e-commerce administration platform supporting inventory management, sales analytics, financial reporting, and customer operations.

  1. Built RESTful APIs powering real-time dashboards, reporting modules, and inventory management.
  2. Implemented secure authentication, role-based authorization, and interactive KPI visualizations.
  3. Improved operational efficiency by centralizing product, order, inventory, and sales management.
aldyth.ai Brain
Email

Ask anything about Aldyth

Use the portfolio brain to answer questions about experience, projects, tech stack, referrals, or what he can help build.

Ask about the profile

Search across work history, projects, skills, and opportunity links.

Ideas for writing

Hiring
Portfolio knowledge base