DWARF5 is a new debugging standard that incorporates GNU extensions to better map binary constructs created by optimizing compilers back to original source code while reducing debugging information size. The standard uses multiple strategies including separate .debug files, .gnu_debuglink, build-ids, compressed ELF sections, debug types in ELF COMDAT sections, DWARF Compressor DWZ, .multi files, DWARF Supplementary Object Files, GNU Debug Fission, split-dwarf .dwo files, and DWARF Package Files .dwp files. The basic structure of Debug Information Entries (DIEs) with attributes per compile unit remains fundamentally unchanged from DWARF version 2, but the hierarchy of representation and where bits are stored has become much more complex, making it no longer possible to simply index DWARF descriptions through offsets.
DWARF5 and GNU Extensions: Mapping Binary to Source in Modern Debugging
Added:okay well hi mark whether this my employer I work in the the first tools group with Scouts I hope still with sometimes changed names and I mostly work on fragrant elf you till system tab well in debug stuff so when I wrote the site of going from binary to source that's just part of what dwarf is because if that was all there is then you just need the debug line mapping you take an address out that's the source and wine but dwarf does well so much more you want to know which functions there are what parameters are in the function other variables what is the scope of the variables what are the types of the various things so you have that and then given those variables you you want our values the location descriptors which I thought works pretty well but then well you saw the last talk you want to know ranges of things so you've debug ranges you want to know how you got in a function so drive provides you unwind information not just to know where you came from but also to to get back at the calling wreck the values of the registers when a function was called you can inspect the values at the previous frame it even provides although most most compilers don't emitted directly macro so that if you ever you know the source you want to copy place it you can use the macro to expand all the defines in your code okay that's a lot of things and dwarf has a lot of design goals probably a bit too many because they conflict with each other of course it's kind of interesting that the last talk was a lot about well that's defined by the ABI because one of the design goals is that its implementation independent so in principle for example unwinding you can do almost without any ABI knowledge the the only thing you need to know is which register number corresponds to which actual register on that architecture so it's kind of interesting to see that it's not always done but what is really nice at least I think so now I talked to Tom and some that's the worse part that's not not really true but various places dwarf defines vendor extensions so to reserve the space of constants if you use that then go ahead and add some extra over or forms and I actually think that works out quite nicely because a lot of new vendor extensions that then get standardized later on of course the problem with the fender expenses is you have to implement them in all the producers and all the the consumers for them to actually work and I think Tom's issue is that then when they get standardized you get to implement them all again but but I think it it works quite well this is the one of the design goals that I picked out because I had to pick something to talk about and when I started to list everything that dwarf I've got from the various new extensions it was way too many it's this design goal you you tools that have to process 12 don't have to know about Worf and I think that really helped 12 become quickly usable because it means you can easily compose pieces of code with dwarf and combine them of course this bites some of the other design goals but I thought there were some some clever extensions integrated into the r5 that counter some of these limitations yeah go look there for the whole standard and then we we just discuss a couple of new data representations so to compose to our form two objects which contain dwarf you you don't need that much so this you have an assembler where you can reference labels in the data section you want a way to reference between sections so for example you have the info tree which describes a variable and the test to reference we are the location descriptor is in the other data section you actually have addresses of symbols which you don't know yet where they will be so you you need some way for tools to do the relocation and it would be nice if the assembler knows how to to create lab one to eight I don't know how you pronounce it but a a compact form of writing are constants because otherwise you have to do that and it's a pain but that's that's it with with just that which is basically a generic assembler linker you can combine twelve one of the nice things about that is that you can combine different twelve producers if you first see for the assembler file you get two object files and they get combined even though they produce they have different dwarf producers you don't have to know anything about the dwarf they get combined so to do this show of course not you know now it's a tool well it's somewhat readable it's I wanted to have an example where it wasn't completely trivial but even this actually produces too much to us and sadly it also doesn't show size reduction because on the other hand it's way too small so the overheads just dwarfs it so all the files are almost the same size even though I want to show reduction inside anyway what we what we have is a header file with a simple struct and defines a function that takes a pointer to set the struct and we have two files yeah this will most likely crash and burn so this defines the the function and you have an main depth cost function and the nice thing is that you just create 200 files and you combine them together and the linker doesn't need to know about the different life what the broad data precisely represents so we just looked at the debug info for this so you have the first compiled unit which comes from F dot C the flop function errors name was resolved and the statement the statement list is the line table to use for this and that's at offset zero and that's in principle all the relocation you need in in this case the matching debug line was placed at start because the first object file that you will see for the that for the second one that the statement list is at the next of offset and all the link has to do is place the debug line pieces after each other so in the the thing you see immediately immediately here is in this in both the first and the second compile unit we just redefine that that that such a type which is a bit wasteful especially given that one of the jungles of the earth was that doesn't shouldn't repeat itself too much but of course this comes from we want as simple as possible tools to combine drive in object files yeah let's go back precisely that's the conclusion yes so it's a simple in captivation if I loose it well what happens if you combine these options to doctor the link eliminates the it's it's in that's a nice question because the dwarf isn't touched so if the dwarf in the original objects files describes the function that you then illumine eliminate later it's in there yeah so and indeed one of the questions is if this design goal is but what so what if we put all the all those types in the their own section and what if we add tools that had a link once section or section group which is a good question to ask because that's a from that sections and link are supported so that's a nice thing then if we would calculate some hairs or checksum over the type that we put in in a special section that we could reference those sections with the network SS name and well that's basically what deeper types do it was actually an extension for dwarf 3 in 2012 for and the only thing you then need is a way to reference a a type with with a signature or hair so the the problem is that it does need a new 12 unit header so you couldn't easily combine them and for that reason in others they were put in a different section so wrote note first example nope yeah yes so not so far so DC implements this with Dybbuk type sections we compile it again and this time we see the same debug info let's see if I can find flop and flop C F F is defined as a type 2 D where ok bears which is no no I have to one time I knew I made my example to be sorry now I've run not so okay the argument F is pipe 92 it's a pointer type and pointer tool oh wow that one and the nice thing is that in the other compiled unit we should have the same when we posit yes when we when we create the variable it has the same signature I'm sure that's the same one and now we have a debug type section with that with that same signature and there is only one of them there is yeah there's so that that's really nice because we need another linker mechanism but now we D duplicate some information for free although as my example I was trying to show that it really saves memory and actually they're also object files are slightly bigger because this one structure is too small for the header so it doesn't really help there but you most programs includes larger here there's more types what some type so it actually really helps their the the problem with it is that it makes things more complicated so now to buy data sections it's not enabled by default in GC even though I believe don't try to make it the default sorry well one of the problems was that everybody was kind of lazy so most consumers have offsets into the data section and they just keep an offset and now they have to keep an offset and oh it could be one of those sections I must say I am I'm I'm showing the new extensions but in the par 5 what they did was extend the the die the compiled unit headers so you can now combine five units and all kinds of different units in the same debug info section so that's actually simpler but they also added more complexity that yeah so it's it's it's still very large and it if you look in the example you see lots of relocations references to other data sections and the linker actually has to deal with all that so what what could we we we need them really forced for Strings it's really nice if the tools know how to merge strings because a lot of our data is strange and you really want those to be merged and of course for the other symbols you need relocations and you use relocations for the intersection references and what if we could do something about that and as always the answer is you add a layer of indirection of course so what we do is we put all the addresses in their own section and that's basically just an index and then you can use an index into the address a debug address section and the same we do with strings we add a new assets section so you have instead of having to point directly to a string and asking the link fix this up if you put the string somewhere else you have a new of an indirection so you can just I want string 1 and just relocate the addresses of strings in that section and then finally instead of for example the location the scripture descriptions saying okay my yeah for tracks instead of saying my location the scripts is at that offset in the debug log section or the range you want to range what we do is we add a an index at the start of the section and we at the start of the compiled unit we just say I'm using the debug log section that one and from there you just again use in the XS to where your real ranges or what they start one of the nice things is that true about if we use all those in directions we can move most of it out of the objects valves so we don't need a linker to even see the most of the dwarf data it still has to see the addresses but even the strings can be moved away and this is kind of interesting because so I first show how see yeah let's so what we do here is this is actually the the wharf for new extension that kind of the same and I should have shown that it now creates about an an oval and a DW all twelve object 112 of pixel and a drop objects well it's well but what you see is in the object file you you instead of a real compiled unit here for skeleton and well there are some addresses there and it's there are some section offsets it is the bay at the base address to index into the address list and at F say hey another signature and a a name and well that's it and for the other compile unit it's the same well different ideas different DW all files and this is really nice except if you want to look at the debug info because now the the linker doesn't have to deal with most of the dwarfs data so for 4l few tools with Delphi I implemented info Plus which picks the the after the skeleton it shows it's too technical I thought I it that's lean should actually show from which how it cut now it doesn't and you you see in the DWO file it it it it uses indirect string references and let's see what okay well and it can hmm this wasn't supposed to happen this address comes from from matching up with the the the skeleton seeing where the debug address is and then using an index into that which doesn't work why not and it doesn't every time too bad sorry at least it's consistent yeah sign and is also a small difference with one five year it it still says it use the section offset it doesn't really it should use a different form in into r5 but the idea is it uses the section offsets from this DW all file ten minutes okay oh man that's yet so the same for for the other skeleton and you look up almost most of the the info in the DW over there so this is really nice for your normal edit compile debug cycle because your linker the compiler has to put the debug info aside for most of it but the linker doesn't need to to see it that need to copy it we do have lots of duplications again even for the strings now that those were already original individual files so 12 consumers need to be a bit smarter they have to do some of the deduplication that the linker would do it's awful for this solution because some programs are links of thousands of our files and you really don't want to ship thousands of DW of us I am also not sure people have used this in anger yet really weighed thousands of the wfl's because for a few tools I try to be smart and it's all very lazy so we just open the dwrl file so now you read them when we need to and after a couple thousand you don't get more for the spectres that's not nice I ask you gdb got something smarter no so because that is the kind of concern if you're not in your edit debug cycle being you toast comes with the dwarf tickets well that twelve picketer i don't know how it kind of does what you expect to do it it reads the drafts data it sees all the references to the dwl files and adds them together it acts as a kind of mini linker because strange sections make up so much that you really do want to merge all the all the strings but it's nice because you only have to update the offset table and because you concatenate data sections luckily you only have that one relocation against it what it does is it it creates a index section therefore every dwo it says for the info part and ever Parton well all the debug sections where it starts and ends and in principle that's all you need because you have a lot of indexes that you now know start it well zero or dead offset in the concatenated fourth stuff so of course the next kind of logical step is we're getting to a point where the tools aren't simple anymore so why not write a tool that understands dwarf and does a lot more and that's basically what TW set does it the nice thing about if you know about dwarf and you don't mind one extra extension which is actually also in power five now is that you can be duplicate between the work files you need new forms again to to follow string pointers or references but now you can say oh you for all those debug files just look there I've made one large string section or info section pipes so five minutes I wanted to go in more detail but the number of references between the various dwarf data sections is kind of daunting especially if you'll then combine it with split worth and I thought it was kind of funny about these come from the the dwarf spec that they didn't even try to have split dwarf + supplemental files because I thought I don't think anybody does that yet but it is all specified as you should be able to have split types a split wire and supplemental files together alright yeah so I just showed that there are still backs but I implemented most of this at least the new extensions with an eye on how they are also done done slightly differently in into r5 so if you are writing a dwarf consumer right maybe you want to try out let DW and write documentation for its library and every time I say could you do you want to use it people say okay where's the documentation yeah well there's a header file there's a header file yeah no but it's it's a nice library I'm kind of proud that we keep a B ABI compatibility even while we support newer versions of dwarf of course the value is kind of you'll often need to use new functions to really get at some stuff for example the displaced wife everything works except that both programs not using the newer functions only see the skeletons it's it's nice yes did I do it right sometimes ok questions yes it does need to and there is actually another extension but that's an extension to elf the Alpha format where the compiler compresses all the the dwarf Dayton and the linker expends concatenates and compresses it again there are people will say that's useful oh yeah yeah you can also do that yes and you you can also use LPL's elf utils compress which actually does that for you look it's nothing stops our elf confess sorry I don't even know my own tools sorry yeah stupid small screen so by default it compresses all the data sections for the Edit debacle I would recommend Network but that gdb supported and articles now support it simply tools using their photos libraries need to be updated to support it but it's really easy super easy API so but for TV I would but I'm not completely sure people have used it in anger and there are some questions okay in instead of the linker spending a lot of time you know if gdb opening a thousand files but out of time okay
Up Next

How C++ Debuggers Work: Breakpoints, Stepping, and More
@CppCon
30.9K views•2018-10-20

Solving the Heat Equation with DeepXDE and PINNs
@Dr.Mohammad_Samara
8.5K views•2023-07-17

HTTP Requests Explained: GET, POST, PUT, DELETE
@codecademy
103.1K views•2021-10-07

Enigma Machine Mechanics: WWII Encryption Explained
@JaredOwen
13.2M views•2021-12-11
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Computer Science




































