Back to all projects

AI Coding Benchmark

Web Development Task Generation & AI Evaluation Framework (Perfect WebGen 2 Project)

Frontend Benchmark DeveloperAI Data Labeling & Code Expert / AlignerrOct 2025 - Dec 2025

Developed production-like frontend engineering benchmarks used to evaluate large language models on realistic web development tasks and software engineering workflows.

Project Detail

Screenshots coming soon.

This project does not have local screenshots in the project photo folder yet, so the detail page focuses on the case-study notes below.

Case Study Notes

What this project proves.

01

Created reproducible development environments, task specifications, and reference implementations.

02

Automated functional and visual verification using Playwright, Puppeteer, and browser-based testing.

03

Improved benchmark reliability by eliminating ambiguous prompts and strengthening evaluation criteria.

aldyth.ai Brain
Email

Ask anything about Aldyth

Use the portfolio brain to answer questions about experience, projects, tech stack, referrals, or what he can help build.

Ask about the profile

Search across work history, projects, skills, and opportunity links.

Ideas for writing

Hiring
Portfolio knowledge base