Hugging Face released SmolVLA on 3 June 2025: a 450-million-parameter vision-language-action model with training and inference guidance. The work uses open community datasets and includes experiments with SO-100 and SO-101 robots.
Its immediate value is learning. Engineers can practise recording demonstrations, defining instructions and testing new object positions. Educational arms have different durability, precision and safety characteristics from industrial equipment.
Start by sorting three lightweight objects into three containers. Keep training examples separate from evaluation data, then vary lighting and positions. Document failures as carefully as successes. Use what the experiment teaches to define a production hardware assessment, with a clear budget and decision point.
Original sources
Source summaries and application analysis by COCON. Pilot ideas are proposals, not claims of completed COCON projects or confirmed product availability.
Could this work in your operation?
Share your workflow, location and target outcome. Start with a feasibility assessment, then validate the technology and define a pilot.
Discuss your project ↗ Email our team ↗

