وو

وحید آنلاین . آرشیو وبلاگ وحیدمی دات آی آر . شرکت بیان. vahidmy.blog.ir

وو

وحید آنلاین . آرشیو وبلاگ وحیدمی دات آی آر . شرکت بیان. vahidmy.blog.ir

Len


Len  ...



The length of a data set is given by LEN reserved word:


There are two different LENs:


Win32 Structures according. Example:


[HeadLen: len  value1: 0     value2:W$ 0     value3: B$ 0]


D$HeadLen = 11 (not 7) because in Win 32 Structures, 'SizeOf' is part of the data.


Free use. Example:

          

[data1_2_3: 1 2 3      Words_data: W$ 1 2 3      Bytes_data: B$ 0  

 My_data_Length: len]

          

In this example, memory Dword (len is always a dWord) pointed by My_data_Length contains the value: 19 (19 bytes). Once 'len' has been used, it is  zero:

          

[data1_2_3: 1 2 3              First_Len: len

  Words_data: W$ 1 2 3    Second _Len: len

  Bytes_data: B$ 0    Third_Length: len]

                    

D$First_Len = 12; D$Second_len = 6; D$Third_Length = 1

 


Pay attention that these two Ways for using LEN give different results (+/- 4) and that you can't mix them: If you use 'free LEN' after 'header LEN', 'header LEN' would be turned zero ( > error message). You will reserve 'header LEN's for Win32 Structures only and you will certainly prefer 'free LEN' for your own use (much more consistent with assembly programming).


          

When begining a new set of data (a new open bracket), Len is turned zero. Note that 'len' doesn't gives an equate absolute like '$' in older assemblers; it is true data you can modify at run time... the traditional '$' is a 'little use' feature you can replace in RosAsm with:


[Val1: 1 2 3  Val2: 50    Hand_Made_Len: Hand_Made_Len - Val1]

          

 Of course, it works too (D$Hand_Made_Len = 16).



If you want to set Len to 0 'by hand' (for example for some tests on a data set development), you can simply do:


[Datatest: 0 2 4  len   SomeWords: W$ 3 3   Length: len]     ; >>>  D$length = 4


(First 'len' used for reset).

         

~~~~~~~


Len

Virtual Data

    

Virtual Data   ..



'?' sign is reserved for Virtual Uninitialized Data declarations:


            

[MyData: ? #500]


Will declare 500 uninitialized dWords that will not appear as as many zeros inside the dead file and, that is, will save disk space .


You cannot mix Uninitialized data and Initialized data inside the same Brackets set. RosAsm does 2 different computations for these 2 separate kinds of Data and 'pastes' this chunk after initialized data computations.


 Virtual Data are virtually stored in write order, at the end of Data section. This way for doing this preserves from 01000 hexa memory alignment of an additional .BSS section. Alignment rules, between sets, are the same as the ones for other Data.


As opposed to previous versions, these Data are now zeroed at load time. You have no longer to initialize them at run time, as I previously recommended when zeros were expected.


Beginners: Of course, you will find it much easier to reserve memory for your tables this way, because Initialized and Uninitialized Data accesses are easier than Dynamic ones. For important memory requirements you MUST learn how to ask Win for dynamically allocated memory. Have a search in RosAsm source for 'VirtualAlloc' / 'VirtualFree'... which are the simpliest regular functions for memory management.


When to use '?' Data


Declaring one single dWord as uninitialized Data doesn't make any sense as there are very few chances this really decreases the file size, because of sections alignments. As well declaring it as '0' common Data.


Tables bigger than 01000 hexa are better declared as dynamic memory. Uninitialized Data are perfectly designed for tables, let us say, between 10 and 4000 octets. Typically, several hundreds '?' bytes tables will save memory consumption compared to 'VirtualAlloc' / 'VirtualFree', if we need them running , because several of these tables will be supposedly fit inside a same 01000 hexa memory chunk at run time.


If you are in need of a data table that should be available, as is, from PE upload to PE close, you can do it this way, what ever size it be (in the real world, this is quite uncommon).


So, as a whole, you have a lot of ways for memory management under Win32:


- Data

- Virtual Data

- Global/LocalAlloc

- VirtualAlloc


.... depending on each case... I never use Global/LocalAlloc because it requires a handle plus  pointer management, plus calls to Global/localLock/Unlock. The advantage of these last functions, compared to the 'VirtualFunctions' simplicity is that they do not align on section boundary, and so forth preserve from spoiling memory. (I do not care of Memory spoiling in this case because I need memory either temporary or not in my own writings...).

~~~~~~~




Data Access


Data Access  ....



Simplified access to label and data (reformed syntax):

        

When you want to declare data, you first declare a label (any non reserved word followed by a colon.


         When you want to access the value at a label you give the size marker.

         Two signs are available: '§' (paragraph) and '$' (dollar); 

         (B$, W$, D$,Q$  //  B§, W§, D§, Q§ ), before the label name; That's all.

          

[My_Value: 0]

          

mov eax d§ My_Value   ; d§ / d$ > Dword,    w§ / w$ > Word,    b§ / b$ > byte

 (Same:

  mov eax d§_My_Value

  mov D§MyValue, ecx)



I provided 2 different signs '§' and '$' to make it as easy as possible on all national keyboards. Just choose the easiest one for you.


          

 Sizes of data may be specified in data declarations. Default is Dword:

          

[data1_2_3: 1 2 3     Words_data: w§ 1 2 3     Byte_data: b§ 0]

          

In this example of three labels, the first three data are Dwords. The next three are Words and the last one is a byte.



In MASM / TASM / A86 syntax, data symbols can be declared either as label, or as value:


          >  MyFirstDword: DD 24    ; with colon

          >  MySecondDword DD 25    ; without colon


and can be accessed, in both cases, either as labels or as values:


          > mov eax dWord Ptr[MyFirstDword]  ; the value

          > mov esi MyFirstDword             ; the address

          > mov eax MySecondDword            ; the value

          > mov esi ADR MySecondDword        ; the address (or OFFSET)


plus, of course:


          > lea esi MySecondDword            ; the address


The result of these 'powerful features' is nothing but  utter confusion between values and addresses at writing time.


RosAsm and NASM are much more  consistent: In code reality, values can't be reached anyway out of  addresses . So, any data symbol is an address and can be nothing else. The fact that, we more often use values than addresses in our sources, is not a counter-argument. 


Any attempt to reverse the good syntax, produces defacto the same results as MASM and A86, because of some  particular encoding side-problems (in some cases the assembler can't guess the size, and so, we have to tell it >>> and so on...)


So we have to tell the size at each data evocation. As having to write so many 'Byte Ptr[My_Value]', or even the less difficult 'B[My_Value]', could be boring, I have chosen , to break with the usual convention for 'B§My_Value'. I have implemented '§' and '$'  (paragraph / dollar) as size  markers in order to make it easy on all national keybords.



For Selectors Override, the syntax is, for example:


> mov GS:W§edx-4 00_1101



For Memory Direct Access, the syntax is, for example:


> mov D$+0401130 35


With the '+' Sign, because this is nothing but a Displacement at the encoding point of view. As this formulation is  does not conform to Win32 programming, I did not write a special case Analysis to allow :


>  mov D§040110 35 


(which should usually be a typing error).


~~~~~~~



Jumps Sizes


Jumps Sizes .


         

The Jumps in a Code Section can be either long (four Bytes) or short (one Byte) signed Displacements.


The RosAsm Assembler does not optimize the Jumps Sizes. That is, if you declare a long jmp where it could as well be short, it will remain long.


This optimization that some Assemblers do, has no effect on the speed of the Code. It only concerns the size of Code. As short Jumps are limited to -128 Bytes upward and +127 Bytes downward, the real interest of these shorter forms is also limited to a very local scope.


There are two x86 Instructions that can only be short, and never long: LOOP and JECXZ. Inside a loop, saving three Bytes here, three Bytes there, may effectively make a real difference, as it may become an organization problem, when the Block of Instructions comes closer to the 128 Byte limit.


This is one of the reasons why the RosAsm Assembler implements the Local Labels with Sizes and Directions Markers: Locally, you take care of your jumps sizes. In case of an overpassed limit, the Assembler error Message tells you the number of overflowing Bytes, and then, its on you to take the decision to make your loop shorter, with other Instructions, or to re-organize the given Chunk of Code in another manner.


Another reason is that this choice between short and long forms of the jumps enables the RosAsm Programmer with an easy and friendly method for preventing a very usual developement error: The missing Label Declaration.


Let us take an example with the standard If Macros set, with the leading Periods, for the various If levels. You should always start the development of an If case with the short form. That way, in case you would forget to write the required End_If, the Assembler will point the error out, because, even if some End_If is to be found somewhere else, downward, the distance will much likely be too long. Otherwise, this type of error could be difficult to debug.


~~~~~~~


Local Labels

     

Local Labels  ...



Note for the Tree View feature: if you use plain labels inside your routines, tree view will not give you a very good image of your source structure (each plain label is considered by tree view as a routine entry point). If you do not want 'meaningless' labels in your sources, use HLL-style macros (IF / DO /,  and so on).

     

Inside routines, meaningful labels' names are of little use and create more problems (how to find so  many new names) than they would increase readability:

         

jne L0<               ; is just as clear as:


jne tryAgain


The best is:

 

While eax a ebx

    ; ...

End_While


('While' / 'End_While' defined as user macros including local labels -beginners: You will find Assembly much easier to learn with HLL Macros, but try to use Local Labels from time to time, so that the HLL-style default Macros I propose do not become a mask between you and the code reality-).



A local label is a CODE location symbol you can use as many times as you want: one letter followed by '0...9' followed by a colon:

         

A2:

X7: 

         

are local labels. Yes, same as A86 ones, but evocations of local labels can have a direction and size specifier in all cases:


L3:

...

jc  L3>>    ; long conditionnal jmp to following L3 label

jmp L3<    ; short jmp to upper L3 label

....

....

L3:

         

At  writing time, I recommend that you set all evocations short (< or >): If long specifier (>> or <<)  is needed, RosAsm will tell you. Having to set it long is a good indication of too long unstructured constructions in most cases. 


More: 


This is sometimes a good error signal for a bad branching attempting to reach a label in another routine after some text handling mistake (RosAsm does NOT control local moves crossing over plain labels, so that you can mix local and plain labels if you wish to -you are allowed to do it dirty-).


         

If you write 'jmp L0', of course, RosAsm will suppose it to be 'jmp L0<' (default is 'up-short')


When using these meaningless Local Labels, after some time spent at modifying and re-writing a chunk of code, your Labels order may be turned confusing. As re-ordering properly many Local Labels Declarations and evocations may be painful and a good source of errors, the Source_Editor (see there, the Re-Ordering of Local Labels paragraph) has a friendly feature able to do the job for you.



I recommend you reserve 'L.' for common code and 'M.' for use inside macro definitions. One letter (10 local labels) is more than enough for code; if not, this is a good indication that source needs a more structured rewrite. I often reserve 'L9:' for 'CaseOut' and 'L8: / L7:' for 'CaseError'.


Another good rule is to reserve specific local labels letters for some macros which produced labels could conflict. One example: if you like high level style conditionals (select_case, if, on, ...) you will need  ability of mixing these macros in order to handle any complex case:



[Select_Val: 0]


; association of 'C' letter with 'Case':


[Select_Case | push #1 | pop D§Select_Val | jmp C2>]

[Case | Jmp C9>> | C2: | cmp D§Select_Val #2 | j#1 C1> | jmp C2> | C1:]

[case_Else | Jmp C9>> | C2: ]

[End_Select | C2: | C9: ]



Select_Case eax

           Case e 0

             ; ...

             ; some code lines

             ; ...

           Case e 1

             ; ...

             ; some more code lines

             ; ...

           Case_Else

             ; ...

             ; some more code lines

             ; ...

End_Select

         

Will be expanded as >>>

(If we want to be able to handle negative conditions, this is the usual writing; see more accurate formulation in Double_Negations).


push eax | pop D§Select_Val | jmp C2>>

           jmp C9>>                              ; this first one is dummy

C2:      cmp D§Select_Val 0 | je C1>

           jmp C2>

C1:        ; ...

             ; some code lines

             ; ...

           jmp C9>>

C2:      cmp D§Select_Val 1 | je C1>

           jmp C2>

C1:        ; ...

             ; some other code lines

             ; ...

            jmp C9>>


         ; ......


C2:      ; Case_Else code

C2:

C9:      ; End_Select


  

But you will not be able to nest some more 'Select_Case' inside a 'Case' (we are not in Power Basic with one simple macro...): Labels branching would conflict. So, you will be in need of adding some other macros to handle the other levels (If Else_If Else End_If  /  single-line If, Do Loop_Until, and so on). This will do without any branching conflict if you reserve one label letter to each macro construction ('I' for '.if', 'O' for 'on', and so on). These macros constructions are really great for increasing readability and what little performance is lost is, in most cases, out of purpose.


Another very good solution to this problem is to write, for example, as many macros sets as levels needed with additions of leading '.':


[.If ; .....  ]  [..If ; .....  ]  [...If  ; .....   ]    ; (highly readable...)



Again, local Labels are for Code only. You cannot use them for Data (absolutely impossible anyway). Meaningless Labels in Data may be provided by Macros Data declarations with the reserced '&0' internal Labels management of Macro Parser.


~~~~~~~


Plain vs Local Labels


Plain vs Local Labels  ...



From the point of view of encoded jumps sizes, RosAsm is a MonoPass Assembler. This means that the Encoder does not perform any kind of computation or optimization for short or long sizes of any jmp Instruction. For example, in a MultiPasses Assembler, if a jump can be short, and that the reserved space for encoding the Displacement is 4 Bytes, another computation will be run in order to adjust the reserved space to 1 Byte.


RosAsm always encodes Plain Labels as long Displacements (4 Bytes). Only Local_Labels Displacements can be encoded Short (1 Byte), if the Source Instruction has no direction Marker (default is Up-Short) or if it has a Short Marker (< for Up, or > for Down).


The inconvenience of this technical choice is that, if the user does not control by himself the code Sizes and makes use of Plain Labels inside the Routines for short moves, the produced Code will not be optimized from a size point of view.


The first advantage is, of course, the compilation speed. Another advantage is that, as this size choice is not done by the Assembler, the user has full control upon the sizes.


Inside Routines, you should never use Plain Labels. Local Labels should be preferred, or, even better, HLL Macros making use of Local Labels, with the 2 forms for short and long. In many circumstances, you will see that the use of short Local Labels is a very good error control for yourself. Avoiding long jumps inside a Routine is a pretty good security, for the prevention of errors like a missing label (inside the Routines), but existing label (outside, in another routine).


This last point is very important: Enabling an error case on bad short jumps, instead of having the jumps automatically optimized by the Assembler provides a earlier detection of a type of bug that may be pretty difficult to point out once running.


As a consequence of this encoding strategy, using Plain Labels is completely forbidden for some particular Instructions that work only in Short scope:


loop Instruction can only be Short-Up and the target can only be a Local Label.


jcxz / jecxz Instructions targets can only be Short scope Local Labels


Otherwise, an error Message is delivered.


~~~~~~~




Local symbols


Local symbols  ....



Local symbols are available for Data Labels, Equates names and Code Labels. They begin with a leading ''@''. When encountering, for example ''@Var'', RosAsm replaces this symbol by:


'PreviousPlainLabel@Var':



[MainDataLabel: @Var1: 12   @Var2: 66] ; Data

 ...                                    ; No plain Label between !!!!!!!

mov eax D@Var2                          ; eax = 66

...

[AnotherOne: @Var1: 24   @Var2: 77]

...                                    ; No plain Label between !!!!!!!

mov eax D@Var2                          ; eax = 77

   

MyRoutine:

[@Var1 2   @Var2 5]                ; Equates

 ...

mov eax @Var2                      ; eax = 5   same as:

mov eax MyRoutine_@Var2            ; eax = 5

mov eax D§MainDataLabel_@Var2      ; eax = 66

 ...

@Label:                                ; Label

 ....

jmp @Label



Note: You should refrain from using Local Symbols for Code Labels. For this, you have Local_Labels, that are more appropriate, and HLL Constructs, that make the Sources more readable. The Right-Click feature, and the Tree-View will most often fail at pointing these Local Symbols correctly.


When  addressing the Value of a Local symbol, you can as well write:

 'mov eax D$@Var1', or easier 'mov eax D@Var1'


The Local scope is between two main Labels.



Example of Win Structure notation:


MyRoutine:


[RECT:   @left: D$ 0   @top: D$ 0   @right: D$ 0   @bottom: D$ 0]

        mov D§@Left  10    ; or the short form:

        mov D@left  10       ; same



But:


[RECT:  @left: D§ 0   @top: D§ 0   @right: D§ 0   @bottom: D§ 0]


MyRoutine:


    mov D$Rect@Left  10    ; here, ''mov D@left'' would be ''D$MyRoutine@Left''.




Local Symbols inside Macros 


Before V.2.004, it was simply not possible to have Local Symbols inside a Macros Declaration. There is no real use of this, anyway, in usual Programming. This possibility has been made available for an advanced HLL Macros Set, like the one required by the BCX Port Project.


Local Symbols are computed in the earlier computings of the compilation. The ones declared in Macros Declarations are computed after the Macros Job. There are some limitations with this usage of Local Symbols inside a Macro.


Example of mis-use:


[macro | pop D@Local] 


Main: 

  [@Local: ?] 

  [Data: @Local: ?] 


  push D@Local | macro


The first Statement will push D$Data@Local , whereas the second one will pop D$Main@Local , which is, of course not what you could expect it to do, when writing this. The problem, here, comes from the order of the Assembler computations:


First it computes all Local @Symbol into PlainLabel@Symbol. At that time, it does not consider the macros Declarations.


Second, the user Source is re-organized for further computations: The Data are moved into a given area, the Macro and Equates are stored, and removed from the Source, and so on...


Third, after the Macros job, a second pass of the Local Symbols computation is run, in order to assume the Local Symbol unfolded by the Macros. At that time, the above [Data:] Declaration is no longer inside the Source, so that the above push D@Local is becoming D$Main@Local instead of D$Data@Local.


So, it is important to respect the coherency of the Local Symbols usage, and to definitively refrain from using any Plain Label inside Procedures, where you make use of Local Symbols.


~~~~~~~






DB Pseudo-Instruction


DB Pseudo-Instruction   ....




A DB instruction is available to declare Code bytes 'by hand'.  This may be  useful for:


Experiments on coding.


Coding by hand possibly missing or wrong encoding in the  future.


Useful for Disassembler when holding weird code ( cryptic Code, Data in Code,...)


To reserve room in code for writable Code.



The syntax for DB is, for example:


DB  01  025  070


(hexadecimal required).


~~~~~~~



DLLs


DLLs  ...



To declare a Label as an Exported Function, all you have to do is to set an ending double colon to the Label name:


MyDllFunction::



In RosAsm menu, [File] / [Output] options will open first a multi-purpose Dialog that allows the user to define the usual values for PEs: Stack and Heap Min and Max sizes and to define the output extension and type. If you choose [a DLL] option, when closing this first Dialog, RosAsm opens a second one, to define the DLL Flags and the default upload address.


The default Flags are PROCESS_ATTACH / PROCESS_DETACH.


The Default upload Address may vary from 040_0000 to 0_8000_0000. 040_0000 is the default upload Address for PE Applications. 0_8000_0000 is the beginning of the space reserved for the System DLLs. So, these two limits should not be used for users DLLs. 


In order to avoid relocations, you can define these Addresses for your DLLs.


Beginners, if a relocation is required, everything works without any problem. So, if you do not understand this yet, just leave the RosAsm 0_1000_0000 default address for DLLs and continue, the OS knows what to do. 


The Edit box for this address expects Hexa values. The given value is aligned on 01000 boudaries by RosAsm. The 0_8000_0000 value is rejected, but the 040_0000 value is accepted (useful to make experiments on DLLs' relocations).


~~~~~~~