> Over the next few weeks and months, we will make the following capabilities available
> Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
> We will also release more technical details on the underlying approach.
- Frivolous use of the term World Model.
- Claims 20 seconds of video, shows only jumpcuts.
Coming soon!
The term "world model" as it was once used in model-based RL can now apparently refer to anything as silly as linear regression. Then again, the RL folks probably borrowed the term from behavioral scientists before them. It's probably best to simply accept this :/
A similar thing happened to "object oriented" which has been misused by philosophers and visual artists alike.
It's from epistemology. It is not there to refer to the subjective but to the objective.
> A similar thing happened to "object oriented" which has been misused by philosophers and visual artists alike
For instance?
But then again I heard the downers have always been the first to leave their dung comments here so let's see...
Depending on the members, there certainly is. Put aside the "dismissers", those who have a habit or a hormonal reliance to cast a "meh". Those who objectively assess according to the input that the development of facts provide may bend their "apparent mood" accordingly. This may be more evident here because in brighter times we may be more inclined to post and submit about more idle intellectual beauty ("complications in ancient clocks"), and in darker times it makes sense that we are more focused on the problems.
I see nothing here but them TELLING us how great it is. Not showing us.
Have you watched the 46s video full screen on a monitor and not marvelled at the incredible 4K detail of the FPV motorcycle racing clip?
Pessimistic opinions on the labor market, views about society and politics that border on the dystopian, as well as exaggerated concerns about datacenter environmental impacts, all seem like a poor fit for this website.
I honestly hope they put an unrealistic amount of wilhelm scream into the learning process, just for fun.
I'm confused, videos contain images and audio ...?
Maybe they presume that after a series of "good enough to some" they may be getting near the Real Thing?
> our mission to develop real-world visual intelligence
Visual is mono-modal, isn't it?
- Video and audio generation and editing through APIs and private weight access. (“FLUX 3 Video”)
- Action prediction through selected research and commercial partners, beginning with mimic robotics (“FLUX-mimic and FLUX 3 Action”)
- Image synthesis and editing through APIs and private weight access. (“FLUX 3 Image”)
- Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)