Degree

Honors Baccalaureate

Department

Mechanical Engineering

Document Type

Thesis

Abstract

Industrial automation increasingly demands robots that can adapt to dynamic environments without extensive reprogramming. Traditional robotic systems perform well in structured, repetitive settings but struggle when workspace configurations change. Programming a robot for each new task requires specialized expertise and significant time, creating a barrier for small and medium manufacturers. While recent VLA models have demonstrated strong capabilities in robotics and artificial intelligence (AI) agents, existing research has rarely focused on the complex and sequential manufacturing tasks for robotic handling, printing, sorting and assembling. This research investigates whether CogACT can be fine-tuned to perform a pick-and-place task on the Yahboom DOFBOT without manual reprogramming. A custom simulation environment was built in SimplerEnv, which uses ManiSkill2 as its underlying simulation framework and SAPIEN as its physics engine, and a MoveIt-based data collection pipeline was developed for supervised fine-tuning. Due to GPU memory constraints, fine-tuning was limited to the model’s 160M-parameter action head, with vision and language backbones frozen, reducing trainable parameters from 7.63 billion to 160 million. Because the DOFBOT’s gripper is incompatible with SAPIEN, the physics engine underlying ManiSkill2, a touch-based evaluation framework was implemented measuring task performance through end-effector proximity rather than simulated grasping. Three conditions were evaluated against a MoveIt scripted baseline across five trials each: pretrained CogACT without fine-tuning, V1 fine-tuned on 5 episodes, and V2 fine-tuned on 28 spatially distributed episodes. The pretrained model achieved 0All evaluations were conducted in simulation. Physical deployment on the DOFBOT remains the primary direction for future work, alongside domain randomization for improved sim-to-real transfer, automated demonstration collection, and gripper simulation resolution. The complete pipeline runs in under 20 minutes on a single NVIDIA RTX 4000 Ada GPU, demonstrating that VLA adaptation is achievable within the resource constraints of a small manufacturing environment.

Date

Spring 2026

First Committee Chair

Sen Liu

First Committee Member

Jacob King

Second Committee Member

Boyang Zhang

Share

COinS