Strings in Assembly ..
Under modern OSes, Strings come in two flavors : Bytes Flows and Words Flows. I will describe only the Bytes Strings form here. The same rules apply, of course to the Words Strings form. The Bytes Strings are known as ASCII Strings and the Words Strings as Unicode Strings. Many Api Functions have two forms, ending by ''A'' and by ''W'' that reflects the Character sizes of the Strings concerned by the called Functions.
Unicode Strings are for oriental Languages (requiring many more Characters than the alphabetic system).
So, ASCII Strings are nothing but Flows of Bytes, each Byte representing one Character. Example: the Space Character is represented by the ASCII Value 020 (32d).
A bit of History
The complete Table of ASCII Characters is, - as you may guess when talking of Byte Values - , 256 Characters long. Several aspects of this Table cannot be understood without first considering its historical evolution.
In the earlier days of Personal Computers, that worked in Text Mode, the position of the Caret, on the screen was controlled by the ASCII Characters themselves, in the Printing Interruptions. So, the lower Bytes from 0 to 31 were reserved as control characters. For example, I remember that, at that time, I wrote a Prompt (Command Line Invitation) doing a lot of things at once. It was showing the Path and the Prompt, printing the Time in the upper Right corner, and other things like this, in various colors. The Prompt Command was managing the Caret movements by these Control Characters. Needless to say, these things are obsolete in PC's for a long time, but, the ASCII Codes, for moving the Caret on the screen, for example, still have their place inside the ASCII Table, which was developed with the advent of Teletype machines to send telegrams and telexes worldwide over wires long before any type of computers ever existed on this earth. In fact the the early mainframe computers used Teletypes to allow humans to talk with the CPU.
Linux users should be aware they have TTY devices in their systems. These control characters have their names such as <carriage return> and <line feed> copied from those ancient things called typewriters. Which had carriages, which held the sheet of paper between its platen and rollers, typing on one line at a time. The carriage was then returned to position it to the start of the left column and the platen rolled up to go to the next line, in effect the paper feed. Our keyboards are laid out in the same way as those old typewriters ....in some ways things change and yet remain the same.
There are still in use, even today, mainframes with Teletype interfaces......
You may wonder, for example, why, for going to next Line, we have usually to provide *two* Control Chars: CR/LF (13, 10). Historically, this is because the original meaning, in PC's of CR -Carriage Return- simply was ''Put the Caret in the first Row'', whereas the original meaning of LF -Line Feed- was ''Push the Caret one Line down''.
In the DOS time, the ASCII Table had a version in Memory (Graphics Table), which the upper half of it, was available for Pseudo-Graphical outputs. Usually, this upper part was for drawing Boxes, for example, with double or single lines. Many older Programmers were used to Poking them for building Sprites by Characters. The Good old days... ;)
Organization of the Alphabet inside the ASCII Table
You may view the ASCII Table, from RosAsm, by running [Tools][ASCII Table]. The very first Character to memorize is the Space (020 / 32). Then, two other important positions to consider are the ones of 'A' and 'a' Characters. As you may see, RosAsm's ASCII Table is organized by Row of 32 Characters, that helps viewing the parallelism between the Upper and lower Cases Characters.
For having any upper case Character made lower case, or reverse, all you have to do is add or subtract 32 from its ASCII Value. In fact, we never do this that way. To insure that a Character is Low or Upper Case, we do, for example:
> or B$MyString 020 ; >>> Lower case.
> and cl (not 020) ; >>> Upper case.
Viewing this in Binary (020 = 00_0100000):
7 6 5 4 3 2 1 0
_0| _1| _0| _0| _0| _1| _1| _0| = ''F''
7 6 5 4 3 2 1 0
_0| _1| _1| _0| _0| _1| _1| _0| = ''f''
Another interesting part of the Table is the part with the Numeric Characters. As you see, for translating a Value smaller than 10 into the its corresponding ASCII Character, all you have to do is:
> add al 030 ; If al was 3, it now is ''3''.
... or, same but neater:
> add al '0'
A stupid difficulty comes from this organization: As the 'A' Character does not come right after the '9' Character, (there are seven ASCII Characters between these two), for Printing HexaDecimal forms of a number, we have to add (sub) 7 to (from) the upper values (0A to 0F) because of this 'hole' between '9' and 'A', during the translations between Binary and Hexdecimal.... Too bad... ;)
Strings in Memory
Declaring a String:
> [MyFirstAsciiString: B$ 'Hi! you!', 0]
Most often, Strings are zero ended. For example most of the Api's expect Zero ended Strings: The operations stop when encountering the zero. Even if you are using ASCII Strings, and not Unicode Strings, you may be in need of some Unicode Strings for some Functions (For example, in Resources Templates, the Strings are always Unicode). In these cases, the Declaration is:
> [MyFirstUnicodeString: U$ 'Hi! you!', 0]
Which is identical to:
> [MyFirstHandMadeUnicodeString: B$ 'H', 0, 'i', 0, '!', 0, ' ', 0, 'y', 0, 'o', 0, 'u', 0, '!', 0, ', 0, 0]
Strings in Registers
Any Register may contain ASCII Chars, in the same way they may contain any other Value, with the usual respect of Sizes:
> mov al 'a'
> mov eax 'abcd'
Some older Assemblers, when encoding such a Statement as mov eax 'abcd', perform the Bytes reversing as they do for Values. That is, in Memory, 'abcd', once reversed, is stored as:
064, 063, 062, 061 .... That is: 'd', 'c', 'b', 'a'.
Of course RosAsm does what you expect to have, instead, without reversing these Strings Bytes.
~~~~~~~