Agentic AI BenchmarkJun 2026 - Sep 2026
MINA - Agentic AI Task & Benchmark Development
Designed and implemented production-grade software engineering benchmarks used to evaluate autonomous coding agents across realistic development environments and automated testing pipelines.
- Developed reproducible software engineering tasks spanning backend systems, APIs, databases, debugging, and infrastructure workflows.
- Built automated validation pipelines, reference implementations, and scoring systems for objective benchmark evaluation.
- Improved benchmark reliability by refining task specifications, debugging execution environments, and strengthening evaluation criteria.
01Agentic AI Trace EvaluationMay 2026 - May 2026
Agentic AI Trace Collection & Evaluation (OpenClaw Traces Licensing Project)
Curated and evaluated multi-step agentic AI execution traces used for large language model training, benchmarking, and autonomous agent evaluation.
- Validated execution traces, metadata, artifacts, tool interactions, and outcome annotations.
- Rejected fabricated, incomplete, or low-quality trajectories to improve dataset integrity.
- Strengthened AI training datasets by ensuring authentic planning, debugging, recovery, and tool-use behaviors.
02AI Coding EvaluationMay 2026 - May 2026
AI-Assisted Software Engineering Evaluation (Human Preferences)
Evaluated AI-generated software engineering solutions through production-style development workflows, code reviews, and human preference evaluation to improve large language model coding capabilities.
- Reviewed competing AI-generated implementations for correctness, software architecture, maintainability, testing strategy, and production readiness.
- Performed production-style pull request reviews covering code quality, interface design, error handling, naming conventions, and developer experience.
- Generated structured human preference datasets used to improve LLM reasoning, coding quality, and instruction-following performance through RLHF workflows.
03LLM Reasoning EvaluationMay 2026 - May 2026
Meter Pavilion - AI Evaluation & Reasoning Analysis
Evaluated large language model reasoning, instruction following, and multi-step problem-solving through structured evaluation workflows to improve model alignment and benchmark quality.
- Analyzed complex reasoning chains, edge cases, and instruction-following behavior across diverse benchmark scenarios.
- Produced structured evaluation feedback supporting model alignment, reasoning quality, and benchmark development.
- Improved evaluation consistency through standardized human-in-the-loop workflows for next-generation LLM systems.
04MoodTune AI - AI-Powered Music Recommendation Chatbot
AI Application DeveloperPersonal project
Built an AI-powered music recommendation platform that combines Gemini API and Spotify API to generate personalized playlists from natural-language conversations, moods, and genre preferences.
- Designed an end-to-end conversational AI application with real-time music recommendations.
- Integrated Gemini API with Spotify APIs and embedded playback for seamless user interaction.
- Built a modern Streamlit interface optimized for rapid iteration and AI feature experimentation.
05Agentic AI RL EvaluationApr 2026 - May 2026
OpenClaw RL - Agentic AI Task Generation & Evaluation Framework
Designed, executed, and evaluated complex multi-stage agentic AI tasks in OpenClaw to benchmark LLM planning, tool orchestration, reasoning, safety, and artifact generation for reinforcement learning datasets.
- Developed realistic multi-stage evaluation tasks that benchmarked autonomous AI agents across planning, reasoning, tool coordination, and execution quality.
- Improved benchmark reliability by enforcing reproducible environments, standardized evaluation rubrics, and verifiable artifacts for fair cross-model comparisons.
- Enhanced AI training quality by identifying failures in reasoning, instruction following, safety, factuality, and tool usage for higher-quality reinforcement learning datasets.
- Contributed to robust evaluation pipelines for next-generation autonomous coding and agentic AI systems through structured task design and comprehensive performance assessment.
06Speech Data EvaluationApr 2026 - Apr 2026
Air Traffic Control Audio Transcription & Evaluation (ATC Project)
Validated aviation speech datasets by producing FAA-compliant transcripts for air traffic control communications, supporting speech recognition and language model training.
- Generated accurate aviation transcripts using standardized phraseology and communication protocols.
- Improved transcription quality through speaker identification, callsign normalization, and terminology validation.
- Maintained dataset consistency across noisy audio, incomplete transmissions, and edge-case scenarios.
07Agentic AI BenchmarkMar 2026 - Apr 2026
OpenClaw RL - Agentic AI Evaluation & Benchmarking
Designed and evaluated multi-agent AI workflows using OpenClaw to benchmark reasoning, planning, debugging, tool orchestration, and software engineering capabilities of large language models.
- Created realistic multi-stage benchmark scenarios reflecting production software engineering workflows.
- Standardized evaluation methodologies through reproducible execution environments and structured scoring frameworks.
- Identified planning failures, reasoning limitations, safety issues, and tool-use behaviors to improve future model performance.
- Produced benchmark datasets supporting reinforcement learning, autonomous agents, and LLM evaluation research.
08Agentic AI EvaluationDec 2025 - Jan 2026
Agentic AI Scenario Quality Assurance & Evaluation (Agent Error Analysis Project)
Designed and validated agentic AI evaluation scenarios by reviewing task specifications, execution traces, evaluator logic, and scoring rubrics to improve benchmark reliability for autonomous AI systems.
- Validated agent execution traces against expected behaviors using reproducible evaluation criteria.
- Identified evaluator inconsistencies, ambiguous scenarios, reward-hacking risks, and specification gaps.
- Improved benchmark quality through stronger rubrics, trace validation, and standardized evaluation workflows.
09AI Training DataNov 2025 - Nov 2025
Egocentric Video Data Collection for AI Training (Terminator Project)
Produced and validated high-quality egocentric video datasets for multimodal AI training, ensuring compliance with strict recording, privacy, and annotation standards.
- Collected first-person activity datasets supporting computer vision and multimodal AI training.
- Performed quality validation against recording standards, privacy requirements, and reviewer guidelines.
- Improved dataset consistency through iterative quality reviews and submission refinements.
10AI Coding BenchmarkOct 2025 - Dec 2025
Web Development Task Generation & AI Evaluation Framework (Perfect WebGen 2 Project)
Developed production-like frontend engineering benchmarks used to evaluate large language models on realistic web development tasks and software engineering workflows.
- Created reproducible development environments, task specifications, and reference implementations.
- Automated functional and visual verification using Playwright, Puppeteer, and browser-based testing.
- Improved benchmark reliability by eliminating ambiguous prompts and strengthening evaluation criteria.
11ETL AutomationMay 2025 - May 2025
Data Automation Report Cust Perisai Plus
Built an automated ETL pipeline that centralized operational reporting data for Cust Perisai Plus, improving reporting efficiency and data consistency.
- Designed and automated end-to-end ETL workflows for operational reporting.
- Reduced repetitive manual processing through scheduled data automation.
- Improved reporting consistency and accelerated access to business insights.
12Auto Report SV MCC (Sales Volume Merchant Category Code)
Developed an automated reporting service that generates Merchant Category Code (MCC) sales analytics from large-scale credit card transaction data.
- Built scheduled Go-based ETL jobs to automate transaction aggregation by Merchant Category Code.
- Reduced manual report generation while improving reporting consistency and data accuracy.
- Delivered timely sales analytics to support business monitoring and portfolio management.
13Financial Data AutomationApr 2025 - Apr 2025
Data Automation ENR
Developed an automated ETL service for Estimated Net Revenue (ENR) reporting, transforming financial data into centralized analytics datasets.
- Designed automated pipelines for collecting, transforming, and loading operational financial data.
- Improved report availability while eliminating repetitive manual processing.
- Implemented scheduled Go services to increase reporting reliability and consistency.
14Transaction Reporting AutomationApr 2025 - Apr 2025
Data Automation JCB Report
Developed a transaction reporting automation service that centralized JCB credit card transaction data into enterprise reporting systems.
- Built automated ETL workflows for JCB transaction reporting.
- Reduced manual reporting effort while improving reporting accuracy and consistency.
- Delivered standardized transaction datasets for downstream analytics and business reporting.
15Cust Activation Campaign & Customer Journey Monitoring
Built a customer journey analytics platform that monitors the complete lifecycle of credit cards from issuance through activation and transaction activity.
- Developed end-to-end customer journey tracking from card issuance to first and last transactions.
- Automated backend data processing using Python to synchronize customer activity across banking systems.
- Integrated multiple internal services through REST APIs to support personalized activation campaigns and business analytics.
16Billing App Fintech
Developed a fintech billing and reconciliation platform that tracks partner billing activities, validates payment reports against transaction data, and improves financial reporting accuracy.
- Built billing workflows covering partner usage, reconciliation, payment validation, and reporting.
- Automated reconciliation by comparing external billing reports with internal transaction records.
- Reduced manual verification effort while improving billing accuracy and operational efficiency.
17Database Marketing Analytics
Developed and maintained an enterprise analytics platform that consolidated customer, transaction, and financial data into interactive dashboards for marketing and portfolio management.
- Built KPI dashboards covering customer acquisition, transaction volume, ENR, installments, and merchant analytics.
- Integrated multiple enterprise data sources into a centralized analytics platform.
- Enabled business users to make faster data-driven marketing and financial decisions through interactive reporting.
18Airport Security OperationsFeb 2021 - Apr 2021
e-Bappi
Developed a web-based operational management system that digitized airport security workflows for prohibited, entrusted, and lost property management.
- Built a full-stack application replacing manual paper-based operational processes.
- Implemented secure item registration, tracking, and reporting with real-time status updates.
- Improved operational efficiency by centralizing airport security records into a searchable system.
19E-Commerce OperationsSep 2020 - Dec 2020
e-Commerce Admin Dashboard
Developed a responsive e-commerce administration platform supporting inventory management, sales analytics, financial reporting, and customer operations.
- Built RESTful APIs powering real-time dashboards, reporting modules, and inventory management.
- Implemented secure authentication, role-based authorization, and interactive KPI visualizations.
- Improved operational efficiency by centralizing product, order, inventory, and sales management.
20