The other thing is MSDF rendering is fairly cheap and can be done on the CPU quite easily without GPU shaders, the atlas generation/upload is the expensive part, but it's a one-time cost per character per font as the atlas is just a normal bitmap texture file, MSDF text have sharp edges at most zoom resolutions, and the small size text is better handled with simple CPU raster anyways.
I really failed to see significant benefit of using Slug over MSDF + raster fallback for small fonts, it's definitely more exact, but I'm not sure if the marginal resolution benefit is worth it over much more complicated GPU dependent rendering, so I'd really want to test it out myself when I have the time over taking the word of an obviously AI written article for it.
At first glance, subdividing the Bezier curves sounds like a bad idea (more Beziers to rasterize) but it opens doors for some parallelism, and most Bezier curves that appear in fonts are monotonic in the first place (so the increase is very modest). This was inspired by this entertaining but not very serious video about font rasterization [0].
The first parallelism optimization is checking against the curve bounding box vs. a rectangular (in uv-space) region of pixels, and this can quickly determine if the Bezier needs to be evaluated in the first place. This can be done per GPU warp.
The second optimization works only for rectilinear transformation (no rotation, skew or perspective). Solving the quadratic equation involves a square root and a division (which alone are >30% of the computation), which can be computed for each row and column of pixels instead of for each pixel (2n instead of n^2).
Both optimizations rely on mathematical invariants of monotonicity, ie. the derivative of the Bezier curve must be non-zero. All Bezier curves can be robustly subdivided into monotonic sections using de Casteljau's algorithm.
My simple benchmarks compare favorably to Slug on the GPU and to "fast" rasterization algorithms on the CPU (which is an order of magnitude faster than "fancy" rasterization algorithms with hinting etc).
Unfortunately there are so many hobby projects and so little time. All I have is messy shaders that draw individual characters and a few benchmarks to see how quickly (and something similar for the CPU). Going from there to a complete text rendering system would be a lot of work. Writing a more detailed article with illustrative code examples is something I'd want to do but haven't gotten around to.
If you want to offer words of encouragement or geek out about rasterization algorithms, I welcome any input.
[0] https://www.youtube.com/watch?v=SO83KQuuZvg Sebastian Lague - Coding Adventures: Rendering text.
Slug doesn't seem to support any of these, it is just "for a given point, am I in or am I out?", but it doesn't tell you by how much, which is great on really high resolution displays and large sizes, but you would lose the ability to do the kind of effects you can do with SDFs, and have to deal with antialiasing separately.
It does give you an anti-aliased value between 0 and 1 that estimates how much of a pixel is being covered.
But this is a linear estimate based on horizontal and vertical distance to the Bezier curve. It does not look correct at long distances, which is why you shouldn't use it for outlines, drop shadows or the other cool things you can do with (M)SDF. A single pixel outline works fine but is not really legible with modern display resolutions (very thin lines).
For finding the minimum distance between a quadratic Bezier curve and a point would require solving a 3rd degree polynomial, where Slug's algorithm gets away with solving a quadratic equation per pixel. This is makes a big performance difference.
That said, there are interesting effects other than grayscale antialiasing, and I wonder how well slug could handle things like outlines (useful for subtitles over video, to make them readable on any background).
Not especially, unless you're running something like a high-refresh-rate gaming monitor that's still 1080p or less.
On most current phones, laptops, and monitors, I would expect the difference between grayscale antialiasing and monochrome rendering to be hard to notice.
It does bias me negatively towards the author, and makes me a little sad that people are choosing to not put more thought and effort into their writing, but this is a trend that is here to stay.
Making a big fuss about this hasn't gone anywhere from what I can tell, and it's unlikely that the outcome will be any different going forward.
> There is no atlas, so a hundred thousand CJK glyphs cost a font's worth of outline data, not an atlas the size of a video. Text can change every frame at no baking cost, which is exactly what you want for live data, user input, and localized content. And it all happens in a single draw with an ordinary fragment shader, no vendor extension required.
Specifically, these sentence structures:
1. ... so <blah> costs <blah>, not <blah>.
2. And <blah>, no <blah>.
3. ... fixed resolution ahead of time, which quietly ...
> Drawing it on a GPU, crisply, at any size, under any 3D transform, while the text changes every frame, is not.
> One small texture, resolution independent within reason, one cheap shader.
English is not my native language, so I may be wrong here.
I think it's a mix of human and LLM writing.
It may be of interest to some game/graphics devs here...
the reason slug has two bands is precisely because of its anti-aliasing. you having only one implies you do the AA differently, or are less efficient with searches off the main direction.
for readers: the slug shader traces two rays, one horizontally, and one vertically, to find out if a pixel is inside a glyph or outside. one would be enough for pure inside outside, but for anti-aliasing purposes it's also helpful to know how far away you are from the closest edge. but if you have a horizontal ray running parallel to a horizontal glyph edge (not uncommon), the ray glyph intersection will return no value at all. so slug traces two rays and blends between them for anti-aliasing that always works in both cases. the bands mentioned here are just acceleration structures, basically a list of cells where each tells you which parts of a glyph are contained within it.
it's still approximate. a true ground truth would just supersample and do many point-in-glyph tests within one real pixel - at edges some of them would be inside, and some of them would be outside, giving you a smooth value to display depending on the shape of the glyph within the edge pixel.
“Chinese, Japanese, and Korean have tens of thousands of glyphs, and baking all of them at several sizes is a memory disaster” One possible solution is dynamic atlas built on CPU for visible glyphs only.
“The distance-field panels notch, where interpolating between stored samples no longer matches the true curve” Can’t it be fixed in the shader, using screen-space derivatives of the SDF? I think in theory, SDF value for pixel center combined with screen-space gradient vector of that number delivers enough data to compute partial coverage for the pixels on the edge.
via WebGPU backend: https://floooh.github.io/sokol-webgpu/slug-sapp.html
via WebGL2 backend: https://floooh.github.io/sokol-html5/slug-sapp.html
There's quite a bit of helper code plus stb_truetype.h and stb_ds.h under the hood to parse TTF files and crunch the TTF curve data into the runtime format expected by the Slug shader (this stuff should better go into an offline asset pipeline tool):
https://github.com/floooh/sokol-samples/blob/master/libs/slu...
...the actual text rendering code is also taking a couple of shortcuts, e.g. no kerning, no right-to-left, and also no text shaping.
There's also a new and complete text rendering stack by Mikko Mononen called Skribidi (AFAIK not based on Slug though):
https://github.com/memononen/Skribidi
...the list of external dependencies is a bit scary for a small self-contained sample though (Harfbuzz, SheenBidi, libunibreak, etc...), but that basically shows that proper international text rendering is really damn hard, even when trying to simplify the code as much as possible.
Question though, you implemented the slug system? The article mentions root eligibility but doesn't expand on it, what's the point of it because slug doesn't seem too "magical" for only doing winding counts?
I'm a bit confused because at some points it seems like the author is conflating "tessellation" and "Rive". Are they the same thing? As an uneducated reader, my understanding would be that Rive is an implementation of a renderer using the tessellation approach. But surely a generic tessellation approach could support perfect arbitrary transformations even if Rive doesn't?
Maybe there's something I'm missing.
Also, a nitpick: in the "head to head" section, the author highlights Slug's better performance in green for the entries that it wins (or ties). For the sake of fairness, shouldn't we highlight the winners in every category? Surely Rive's "low" memory usage beats Slug's "moderate"?
It's nice that he gave the patent to public domain, but this is not how patents are supposed to work. You can't patent something two years after it was already published.
I'm guessing he actually filed for a patent before publishing, and the article should read: "Lengyel was granted a patent for it in 2019"
It's an AI-written article, so maybe it's not reasonable to expect it to be consistent on this level...
* The JCGT paper was published a few months later on June 14, 2017, after the priority date.
* The final patent application was submitted on February 1, 2018, before the one-year deadline.
* The patent was granted about 1.5 years later by the USPTO on August 6, 2019.
https://terathon.com/blog/decade-slug.html
The timeline basically checks out, maybe it simply took the patent office over a year to process the patent application?
MSDF is pretty much just the target texel in question plus the surrounding samples in a way the GPU can do entirely upfront before the math starts. (M)SDF glyphs also play nicely with mip mapping and I would think this Slug algorithm needs uncompressed data. Maybe that doesn't matter because you just don't scale the data ever?
The pixel shader is indeed quite complex compared to SDF:
https://github.com/floooh/sokol-samples/blob/8afa83928ce1870...
https://rookandpossum.com/posts/scanline-sweeper/
Sean Barret (creator of the stb public domain libraries) independently invented a CPU-based implementation of the same idea, used in stb_truetype.
The project I'm most known for is basically an MSDF shader with a bloom pass. It serves my needs 100%, though if I expand to support arbitrary text, I may reach for Slug.
The one issue with MSDF I want to raise is, it seems everybody uses the same msdfgen texture creation program from Viktor Chlumský's master's thesis 11 years ago. I wish there were other implementations. Who ever heard of a graphics technique that was only ever programmed once, and then used everywhere without substantial iteration? We need to de-XKCD-2347 MSDFs for everyone's sake, including and especially Chlumský.
https://web.archive.org/web/20120505013814/https://www.valve...
The pdf details only simple SDF, but in the closing paragraphs, it mentions the weakness of the technique and mentions how it can be solved with multiple SDFs, but the exact technique wasn't showcased.
Which led to people trying to reverse engineering it, and making their implementation, for years (it's Valve after all). Including me. Not sure if what I came up with was exactly MSDF, but certainly there are a lot of implementations out there.
I'd say what you and the others did was research, not implementation. Your results were solutions to the same problem, not variations of the same recipe. What you did is important but separate from what's concerned me.
A recipe of Chlumský's— the msdfgen utility— is widely used, but only has one producer. I'm just saying that that's a liability. Like, imagine if HarfBuzz was the only text shaper, and was maintained by one person.
But slug wins in perspective in my eyes.
I use MSDF to render crisp text in my webgl hobby game. Hope to publish it with source code when I get the time.
thanks for sharing the article. I'll take a deeper look at it later.