Tuesday, September 22, 2009

Adventures in virtual texture space, part 6

Yesterday I was working on my virtual texture code, and there's only one bug left (that i know of).

I haven't had much time to look at it, but it fails to load some pages when the page-cache is full for some reason, so it probably has to do with the code that kicks out unused pages when the cache is full.

That aside i'm having some problems with loading my pages from disk.

If I schedule pages to be loaded right when i discover i need them, and just keep adding them to the list, i end up with an ever increasing list if i move around fast enough, because the disk IO simply can't keep up.

So i tried a couple of things, one of them being counting how many times i have a reference to each page in my readback buffer, and giving the pages with the highest ammount of references the highest priority.

After all, the more visible it is, the more imporant it is to load!
That helped somewhat, but still caused a lot of pages to be loaded and uploaded which aren't visible, yet might actually kick out pages that would be more likely to be visible soon.

The next thing i tried was to remove all the pages from my scheduler that aren't visible in the current frame, but that caused pages to not get loaded at all while you move your camera.

A couple of things that i'm going to try:
  • A CPU side disk cache, i only have one on the GPU at the moment.
  • Compression, hoping that'll decrease all these problems simply because loading would be faster and the list wouldn't grow as quickly.
  • Using adjacency information. By pre-calculating a bounding box of all the texel area in world-space, i can sort all the pages by how close they are to each other, and perhaps do some analysis on what is likely to be visible soon, and what would be unlikely to be visible soon.
  • The source material I'm working with is pretty horrible because it has a lot of decal textures which take up extra pages on the screen, and there are simply too many different pages on the screen at the same time. Building my own CSG preprocessor would allow me to optimize this more easily, as apposed to trying to fix a relatively arbitrary list of triangles with all kinds of materials.

My good friend Volker suggested a couple of things i could try:
  • If you know if a page is to be loaded for the current frame, it may be wise to put the higher and lower mip levels of that page in the queue too. (but only if there is nothing else)
  • Kick out pages by random instead of last used or least frequently used caching schemes, apparently the worst case is better, so that's worth a shot.
    (Obviously no pages should be kicked out that's visible in the current frame)
  • Using memory mapped IO.
    Volker: "in my tests, mmap io made a difference from 10mb/s to 500mb/s. (using the disk cache)"
    My biggest problem is latency though.

So does anyone out there have any good ideas on how to determine which pages should be loaded & which pages should be kicked out?

I'm going to see if I can upload a new video of my test level tonight.

Thursday, September 17, 2009

Adventures in virtual texture space, part 5

Today I've spend some time improving my tools and they're much simpler and faster than before, it only takes about 20 seconds to convert a quake4 level into a 16384x16384 virtual texture file and a separate geometry file. I'm quite happy with the tool as it is, there are only a handful of things i still want to do such as trying to rotate an allocated texture space to see if it'll fit better and forcing a texel density to some sane (maximum) limit.

Before i used 128x128 pages, and used texture-arrays (which are basically 3d textures where the z direction alway uses near filtering) with a depth of 512 layers.

This worked out pretty well because you never have any bleeding artifacts between pages.
It is possible to have visible seams between 2 pages that are rendered next to each other if they have a strong enough contrast right on the edge between them, the seam looks as if it's an aliased edge.
However i would consider it extremely rare, i only managed to see such a seam when i purposely made some handmade pages to see if it would happen at all, i couldn't find one when i looked at it in more real-life artwork.

So today i tried using smaller pages, but this caused some problems.
First of all, since i make the pages smaller (say 64x64 or 32x32) it uses less texture space, therefore i need more pages for the same amount of texels.
In theory, smaller pages should be able to match what i render on screen more precisely.
However, since my texture-array has a hard limit of 512 layers, even when the width/height is 4x smaller, i had no choice but to create a version of my texture cache that works with a giant 2d texture.
I haven't bothered to put borders around my pages (yet), so there are plenty of artifacts rendering it like that.

But when i started rendering lots of other artifacts started popping up, which apparently where more likely to happen with smaller pages somehow.
So i fixed a couple of these artifacts, some which helped improved performance, and i still have a couple of mysterious ones left.

Eventually i'll probably build something where i can record and replay a certain path trough my test level and i would use it to compare all the different parameters that can be used to build and render a virtual texture and see which ones are more efficient compared to the other.
This will also help to measure performance improvements, or the reverse, when i'll try to implement stuff like texture compression.

Unfortuneatly i probably won't have too much time working on this in the near future, so i'm kinda unsure if i should start on a CSG preprocessor for the level at this moment, because it'll take a couple of days to build and test.

On the upside I actually managed to get into NVIDIA's "GPU Computing Registered Developer Program"!
Which means i have access to OpenCL, which i would like to experiment with, to see if i can use it to optimize virtual texturing.
I can imagine that determining which pages are currently visible, could be done more efficiently trough OpenCL.
It could be done mostly on the GPU, saving CPU time, and would reduce the amount of data to be downloaded back to the CPU.
Another thing it could help with would be to improve texture decompression speed.

Tuesday, September 15, 2009

Virtual Texturing part 4; importing madness

So after rewriting the code to load the pages in the background on a secondary thread, i started to write some code to import a Quake 4 level (.proc files) and modify the geometry so it could be displayed using a virtual texture, which would be automatically created from the textures in the level.
The red in the screenshot above are textures that i couldn't automatically discover without hacking. The black areas are supposed to be transparent, but i'm not handling that at the moment.

Here's a screenshot where every color is a different page, which shows that i have way too many different pages on screen at the same time:
There are a couple of things that I've learned about converting existing (Quake 4, but the same will apply to other sources as well) geometry to take advantage of virtual textures:
  • Quake 4 uses a z prepass to take care of occlusion, so it's geometry is optimized for number of triangles and not so much for using as little geometry area as possible, which means a lot of wasted texel space.
  • Quake 4 has a lot of transparent textures that are placed upon other textures, which again leads to wasted texel space, as you can see in my screenshots i'm actually not handling transparency.
  • Since Quake 4 has separate geometry for each type of shader, you might end up with lots of patches of geometry that each have completely different pages. If this was build with virtual textures in mind, it would've been continuous. This is bad because it means more pages need to be loaded into memory.
  • Sometimes large textures are assigned to a relatively small area. If you don't take that into account you'll be assigning large areas of texture space to something which is tiny.
  • Without parsing materials (which i'm not doing), discovering the right textures is sometimes impossible.
These problems are causing me some headaches with my test-scene because i'm loading waaaay more pages than i would need to in a scene that would've been build with virtual textures in mind.
I could solve this by building my own quake4 map CSG code, which i might do eventually as i already have some experience with CSG.

However, if you would be building geometry from scratch this would all be easier, as long as you try to keep texel density at a sane level and keep surface area to a minimum. (aka don't assign texel space to something which is never visible)
Sounds rather straightforward, i know, but if you don't think about this up front you might end up with some nasty surprises later on.

Also, allocated texture space should be aligned to page boundaries, if you don't you might end up loading 4 pages when 1 would've been sufficient.

One mistake i made trough this whole process was thinking about a virtual texture as a giant texture, and processing it as such.
The problem with this is that you cannot handle unused pages easily.

Update: scratch that, fixing existing geometry (automatically), to be able to be used with virtual texturing -efficiently-, is hard enough to be considered a dead-end.
I'm going to rebuild the geometry with my own CSG process instead.

Monday, September 7, 2009

Virtual Texturing part 3

My virtual texture implementation now reads pages from disk.

I'm doing it the completely naive way, reading and uploading a page to the video card *just* before i need it, and i'm absolutely surprised how little performance penalty i'm seeing; only 0.1 - 0.2ms!

I'm guessing that this has to do, at least in part, because of disk caching since i'm using the regular .net disk functions at the moment (and i -just- generated my virtual texture disk file, so it's fresh in the cache).
Update: Confirmed, after a reboot i get horrible spikes of +/- 30ms when i move the camera around on the virtual texture!

Unfortunately .net won't get any memory mapped IO until .net 4.0 comes out, so i won't be able to try this unless i port everything over to C++.

I really should 'acquire' some more interesting test scenes.. a flat polygon is simply too ..erhm.. simple.

Wednesday, August 19, 2009

Virtual Texturing part 2

A picture, or well in this case a movie, is worth more than a thousand words:


On top of the screen you can see (pages A/B) how many pages are used (A), and how many there are in memory (B).
The FPS / Ms should really be ignored..

Internally the OpenGL library for C# i'm working with (OpenTK, and i don't have the source-code) or a broken driver is forcing my app to sync to my monitor's refresh rate (60hz), and my recording software is seriously interfering with the timing of that syncing.

Update: putting a timer around the rendering code (and not relying on OpenTK's timer which is always synchronized to the monitors' refresh rate) showed that i was actually rendering around 4000-5000fps (0.2-0.3ms) instead of just 60fps ;)

Other than that there hasn't been any real performance optimalisations yet and i'm doing some things the quick & dirty way (readbacks are not async, page uploads are not async either, and should probably be spread out over multiple frames etc. etc.).
I'm basically just doing whatever i can to get things working like i want instead of trying to use the most optimal OpenGL api mechanisms at the moment.

Next to that there are still some artifacts there, i mean next to the usual video compression artifacts, which i'm trying to track down at the moment.
Just look at 0:30 to 0:32 in the bottom-left corner of the video.
It's weird that the page loading artifacts are there, because at the moment i'm uploading all the required pages before i use them.

The virtual texture is a mere 8192x8192 at the moment and completely in memory (no reading from disk at the moment).
Any larger texture would have to be loaded from disk, my system would probably not like loading a 16384x16384 texture in memory ;)

Filtering is trilinear, Bilinear would obviously be faster and would stress the system less as there are fewer pages to be required to be used per frame.
Seams between pages are more visible with bilinear filtering, but surprisingly hard to spot using trilinear filtering.. in fact, i haven't been able to spot a single one so far!

Next steps would be to fix all the artifacts, start reading back from disk, and optimizing everything.

Friday, August 14, 2009

Virtual Texturing part 1.

Yay! I finally managed to liberate a little time for me to work on virtual texturing!

Thinking it would help me avoid worrying about additional borders i used a texture-array instead of a large texture for my physical page-cache texture.

(Edit: My test textures just happen to be a 'best case' and with alternating border colors between pages an aliased edge is actually visible, so additional page-borders are still necessary)

Each layer in the texture-array (255 max) is a single page, each page is 128x128, which would give me a cache of 255 x 128 x 128 pixels.

Right now the 'virtual texture' is small enough that it fits completely in video memory, so it's not exactly 'virtual', there's no readback yet either.

There's already a page lookup table however.

The next step would be to readback which pages are visible, upload them and update the page lookup table.

Here's a short video showing the blending between the pages

Yes Yes no fancy graphics or even interesting geometry, I'm on a tight time schedule here people! ;)

While working on this i did realize that the rage screenshot in my last post has something odd in it..

Before i only noticed that all the pages where just nice and square and that they had a nice locality to them.

What i failed to notice, and what i notice now, that it's just plain weird that there's just no blending or any transitions between what i assumed where mip-maps!?

Maybe they're showing pages at the highest resolution? And the ones in the back are bigger because some pages just happen to have a larger geometric area assigned to them? And it just, by chance, looks like some sort of weird rough mipmapping?

Friday, August 7, 2009

Rage

I was just reading a new pdf about Rage, which i'm sure you know is the new game from id software, and i noticed something interesting in one of the screenshots:


Notice how the texturing is nicely aligned?

Seems to me that it would be possible to build a rough bounding volume tree with which you could determine which pages are potentially visible.
(Something I'm thinking about with my Deferred Virtual Texture Shading stuff)

Interestingly they're still reading back from the gpu which pages are visible at which LOD.
I wonder if that could be done more efficiently on the cpu avoiding the readback completely.