Tag: Programming

  • chastdin calculator for RISC-V Assembly

    I really did it this time. I translated all of my Intel Assembly functions into RISC-V Assembly and rebuilt my calculator. It works flawlessly running under the rars simulator.

    To run this example, you need the RARS java archive and a java runtime environment installed on your machine.

    You can get RARS here:

    https://github.com/rarsm/rars

    However, once you do, you can run my program with a command like the following.

    java -jar ~/rars.jar main.s
    
    # chastelib test suite for RISC-V Assembly in RARS simulator
    
    # this program tests the stdin extension of chastelib
    
    # The same library of functions I commonly use in my Intel Assembly code
    # have now been translated to RISC-V.
    # All assembly code seen here is for the RARS simulator written in Java.
    
    .data
    
    ##################################################################
    # chastelib core specific variables                              #
    #                                                                #
    # These variables are used by the intstr function to convert an  #
    # integer to a string and what radix and widthshould be used     #
    # width means how many minimum digits including leading zeros    #
    ##################################################################
    
    int_string: .space 32 #reserve space for 32 bytes for up to 32 bits if printed in binary
    int_end: .byte 0 #the terminating zero of the integer string
    radix: .byte 2   #the radix the number will be shown in
    int_width: .byte 1 #by default
    
    # These variables are for outputting special strings
    # such as a newline, space, or a single character based on s0
    
    space: .byte 0x20, 0
    line:  .byte 0x0A, 0
    char:  .byte 0, 0 
    
    ##################################################################
    # chastdin specific variables                                    #
    #                                                                #
    # these variables are used as the default controllers            #
    # for the getstr and getline functions                           #
    # buf stores keyboard input during those functions               #
    # count stores how many bytes were read during system read calls #
    # last_char stores the last character read                       #
    # usually this will be a space, tab, or newline                  #
    ##################################################################
    
    buf: .space 0x100
    count: .word 0
    last_char: .byte 0
    
    # program specific variables
    # These variables are for outputting specific messages
    # or to simulate user input as integers in the strint function
    
    string0: .ascii "calculator for RISC-V Assembly\n"
    string1: .asciz "chastdin (Chastity's STanDard INput) extension\n\n"
    
    string_add: .asciz "add"
    string_sub: .asciz "sub"
    string_mul: .asciz "mul"
    string_div: .asciz "div"
    string_rem: .asciz "rem"
    string_setradix: .asciz "setradix"
    
    string_help: .asciz "help"
    string_exit: .asciz "exit"
    string_putstack: .asciz "?"
    string_clear: .asciz "clear"
    
    string_prompt: .asciz "->"
    
    string_err: .asciz "Error: invalid number or command: "
    string_err1: .asciz "Error: need one number on stack for command: "
    string_err2: .asciz "Error: need two numbers on stack for command: "
    
    chastdin_help: .ascii "chastdin is a stack based interactive calculator\n"
                  .ascii "that reads stdin for numbers and commands.\n"
                  .ascii "Numbers are pushed on the stack for all math.\n"
                  .ascii "Each line can contain multiple numbers or commands.\n\n"
                  .ascii "Arithmetic commands are add,sub,mul,div,rem\n"
                  .ascii "The exit command ends the program\n"
                  .ascii "The ? command prints the entire stack\n"
                  .asciz "The setradix command changes the radix for input and output\n"
    
    .align 2  # Aligns the next item to a 4-byte (2^2) word boundary
    chastack: .space 0x400 #reserve space for RPN calculator stack
    
    .text
    
    la s0, string0
    jal putstr
    
    # change radix for this program
    li t0, 10    #load t0 register with the new radix
    la t1, radix #load t1 register with the address the radix will go to
    sb t0, 0(t1) #save t0 register (byte) to address t1
    
    la s11, chastack #s11 will be used as the virtual stack pointer for this program
    
    #print the help message at the beginning of the program
    la s0, chastdin_help
    jal putstr
    
    #print the initial arrow prompt
    la s0, string_prompt
    jal putstr
    
    main_loop:
    
    la t1, last_char #load address of last_char
    lb t0, 0(t1)     #get the last character
    
    #show the arrow indicating we wait for the user to enter something
    #but only show it when the last character is a newline
    #otherwise it will print too many if multiple commands were entered on the same line
    li t1, 0xA
    bne t0, t1, skip_prompt
    la s0, string_prompt
    jal putstr
    skip_prompt:
    
    jal getstr  # read the string from standard input
    
    #load the length of string just entered from (count)
    la t1, count            #load address of count into t1
    lw t0, 0(t1)            #load number of chars read at (count) address
    beq t0, zero, main_loop #restart main_loop on empty string
    
    #jal putline # print extra line for readability
    #jal putstr # echo it to standard output
    #jal putline
    
    #s0 already contains string that was input
    #s1 will be loaded with address of exit string
    la s1, string_exit
    jal strcmp
    # end program if the string entered is equal to string_exit
    beq t0, zero, exit
    
    la s1, string_putstack
    jal strcmp
    beq t0, zero, command_putstack
    
    la s1, string_clear
    jal strcmp
    beq t0, zero, command_clear
    
    la s1, string_help
    jal strcmp
    beq t0, zero, command_help
    
    #next we begin checking for actual math commands of arithmetic
    
    la s1, string_add
    jal strcmp
    beq t0, zero, command_add
    
    la s1, string_sub
    jal strcmp
    beq t0, zero, command_sub
    
    la s1, string_mul
    jal strcmp
    beq t0, zero, command_mul
    
    la s1, string_div
    jal strcmp
    beq t0, zero, command_div
    
    la s1, string_rem
    jal strcmp
    beq t0, zero, command_rem
    
    la s1, string_setradix
    jal strcmp
    beq t0, zero, command_setradix
    
    
    #if the last string entered was not exit or a math command then
    #The default command is to turn the argument into a number and push to stack
    command_num:
    
    mv s1, s0              #back up this string address to s1 register
    jal strint             #try to get a number from the string pointed to by s0 register
    beq a0, zero, num_push #branch to number push if zero errors in integer string
    
    la s0, string_err    #load error message
    jal putstr           #print error message
    mv s0, s1            #load original command string
    jal putstr           #print which command failed
    jal putline
    j num_push_end       #skip the push because this can't be used
    
    num_push:            #push the number to the fake stack
    addi s11, s11, 4     #increment the pointer by the size of the native int for this mode
    sw s0, 0(s11)        #store the value we converted from the string with strint to this stack space
    num_push_end:
    j main_loop          #once value is pushed, continue the program
    
    exit:
    li a0, 0  #status
    li a7, 93 #exit
    ecall     #environment call
    
    #################################################################################
    # The following functions are used in the calculator program                    #
    # The all jump back to the main_loop after they are done                        #
    #                                                                               #
    #################################################################################
    
    #check if the stack has enough space for the last command
    #this will print an error if less than two numbers were on the stack
    #when using one of the math commands above
    
    memory_check:
    
    la s10, chastack        #load s10 with chastack address for branch comparison
    blt s10, s11, memory_ok # if s10 is less than s11, no errors
    
    print_stack_error:   #otherwise we print error message
    la s0, string_err2   #get error message for less than 2 numbers on stack
    jal putstr           #print error message
    mv s0, s1            #get name of the command used
    jal putstr           #print which command failed
    jal putline
    addi s11, s11, 4     #increment the pointer to what it was before the failed command
    j main_loop          #now go back to main loop after error was printed
    
    memory_ok:
    sw zero, 4(sp)       #if no error, erase the old top of stack by storing zero
    j main_loop          #and continue main_loop as normal
    
    
    command_putstack: #print all numbers on the stack
    la s9, chastack #load s9 with address of chastack
    mv s10, s11     #copy value of s11 to s10
    command_putstack_loop:
    
    #is s10 equal to the address of stack start?
    #if so, end the putstack loop
    beq s9, s10 command_putstack_end
    lw s0, 0(s10) #load the word at s10 into s0 for printing integer 
    addi s10, s10, -4 #subtract the word size from this temp stack index
    jal putint
    jal putline
    j command_putstack_loop
    command_putstack_end:
    j main_loop
    
    
    
    
    command_clear: #erase all numbers on the stack
    la s9, chastack #load s9 with address of chastack
    command_clear_loop:
    
    #is s11 equal to the address of stack start?
    #if so, end the clear loop
    beq s9, s11 command_clear_end
    sw zero, 0(s11) #store zero into the word at 0(s11) to erase it
    addi s11, s11, -4 #subtract the word size from this temp stack index
    j command_clear_loop
    command_clear_end:
    j main_loop
    
    
    command_help:
    la s0, chastdin_help
    jal putstr
    j main_loop
    
    #add number on top of stack to the one below it
    command_add:
    lw t1, 0(s11)     #load the word at this chastack address
    addi s11, s11, -4 #subtract the word size from s11
    lw t0, 0(s11)     #load the word at this chastack address
    add t0, t0, t1    #t0 = t0 + t1
    sw t0, 0(s11)     #save the word at this chastack address
    j memory_check    #check stack for errors after this command
    
    #add number on top of stack to the one below it
    command_sub:
    lw t1, 0(s11)     #load the word at this chastack address
    addi s11, s11, -4 #subtract the word size from s11
    lw t0, 0(s11)     #load the word at this chastack address
    sub t0, t0, t1    #t0 = t0 - t1
    sw t0, 0(s11)     #save the word at this chastack address
    j memory_check    #check stack for errors after this command
    
    #mul number on top of stack to the one below it
    command_mul:
    lw t1, 0(s11)     #load the word at this chastack address
    addi s11, s11, -4 #subtract the word size from s11
    lw t0, 0(s11)     #load the word at this chastack address
    mul t0, t0, t1    #t0 = t0 * t1
    sw t0, 0(s11)     #save the word at this chastack address
    j memory_check    #check stack for errors after this command
    
    #divide and store quotient on stack
    command_div:
    lw t1, 0(s11)     #load the word at this chastack address
    addi s11, s11, -4 #subtract the word size from s11
    lw t0, 0(s11)     #load the word at this chastack address
    divu t0, t0, t1    #t0 = t0 / t1
    sw t0, 0(s11)     #save the word at this chastack address
    j memory_check    #check stack for errors after this command
    
    #divide and store remainder on stack
    command_rem:
    lw t1, 0(s11)     #load the word at this chastack address
    addi s11, s11, -4 #subtract the word size from s11
    lw t0, 0(s11)     #load the word at this chastack address
    remu t0, t0, t1   #t0 = t0 % t1
    sw t0, 0(s11)     #save the word at this chastack address
    j memory_check    #check stack for errors after this command
    
    
    
    #pop top of stack and set the current radix to it
    #it has error checking and leaves the radix as is
    #unless at least one number is on the stack
    command_setradix:
    
    
    la s10, chastack     #load s10 with chastack address for branch comparison
    ble s11, s10, change_radix_no # if s11 is less than or equal to chastack address, branch to radix error
    change_radix_yes:
    lw t0, 0(s11)        #load t0 register with the new radix
    la t1, radix         #load t1 register with the address the radix will go to
    sb t0, 0(t1)         #save t0 register (byte) to address t1
    sw zero, 0(s11)      #erase the old top of stack by storing zero
    addi s11, s11, -4
    j main_loop          #and continue main_loop as normal
    change_radix_no:
    la s0,string_err1    #get error message for less than 1 numbers on stack
    jal putstr           #print error message
    mv s0, s1            #get name of the command used
    jal putstr           #print which command failed
    jal putline
    
    addi s11, s11, 4     #increment the pointer to what it was before the failed command
    j main_loop          #now go back to main loop after error was printed
    
    #################################################################################
    # The following functions are independent of a specific RISC-V Operating System #
    #                                                                               #
    # intstr = convert integer into a string ready for printing                     #
    # putint = prints integer using intstr and the OS specific putstr function      #
    # strint = convert string into an integer                                       #
    #                                                                               #
    # The s0 register is used for pass data in or out of these functions            #
    # See comments above those specific functions for full details                  #
    #################################################################################
    
    # The intstr function does several things at once and is the foundation for all integer output.
    # It uses the global radix variable to know which radix or number base to use when turning the integer to a string
    # It also uses the global int_width variable to determine how many leading zeros should be used for the string
    # The purpose of this is to make numbers look good when lined up when they are printed in a list.
    # radices 2 to 36 are supported. Digits higher than 9 will be capital letters
    
    intstr:
    
    la t1, radix     #load address of radix into t1
    lb t2, 0(t1)     #load value of radix into t2
    la t1, int_width #load address of width into t1
    lb t4, 0(t1)     #load value of int_width into t4
    li t3, 1         #load current number of digits, always 1
    
    la t1, int_end   #t1=address of terminating zero in string
    addi t1, t1, -1  #t1-- to go to lowest digit
    
    digits_start:
    
    remu t0, s0, t2  #t0=remainder of the previous division
    divu s0, s0, t2  #s0=s0/t2 (divide s0 by the radix value in t2)
    
    li t5, 10        #load t5 with 10 because RISC-V does not allow constants for branches
    
    blt t0, t5, decimal_digit
    bge t0, t5, hexadecimal_digit
    
    decimal_digit:   #we go here if it is only a digit 0 to 9
    
    addi t0, t0, 0x30
    
    j save_digit
    
    hexadecimal_digit:
    addi t0, t0, -10
    addi t0, t0, 0x41
    
    save_digit:
    sb t0, 0(t1)     #store byte from t0 at address t1
    beq s0, zero, intstr_end
    addi t1, t1, -1
    addi t3, t3, 1
    j digits_start
    
    intstr_end:
    
    li t0, 0x30
    prefix_zeros:
    bge t3, t4, end_zeros
    addi t1, t1, -1
    sb t0, 0(t1) # store byte from t0 at address t1
    addi t3, t3, 1
    j prefix_zeros
    end_zeros:
    
    mv s0, t1
    
    ret
    
    # this function calls intstr to convert the s0 register into a string
    # then it uses the system specific putstr call to print the string
    # it also uses the stack to save the value of s0 and ra (return address)
    # this way, s0 is restored to the value it had before this function
    # restoring ra is required because it is modified during calls to other functions
    
    putint:
    
    addi sp, sp, -8
    sw ra, 0(sp)
    sw s0, 4(sp)
    
    jal intstr
    jal putstr
    
    lw ra, 0(sp)
    lw s0, 4(sp)
    addi sp, sp, 8
    
    ret
    
    # strint takes the string at address pointed to by s0 register
    # and then loads the s0 register with an integer equivalent value
    # the a0 register is returned with the number of errors that happened
    # programs can use this to find if a user entered a valid number
    # number is intepreted according to the current radix
    
    strint:
    
    li a0, 0         #load zero into register for error counting
    
    la t1, radix     #load address of radix into t1
    lb t2, 0(t1)     #load value of radix into t2
    
    mv t1, s0        #copy string address from s0 to t1
    li s0, 0
    
    read_strint:
    lb t0, 0(t1)
    addi t1, t1, 1
    beq t0, zero, strint_end
    
    #if char is below '0' or above '9', it is outside the range of these and is not a digit
    li t5, 0x30
    blt t0, t5, not_digit
    li t5, 0x39
    blt t5, t0, not_digit
    
    #but if it is a digit, then correct and process the character
    is_digit:
    andi t0, t0, 0xF
    j process_char
    
    not_digit:
    #it isn't a digit, but it could be an alphabet character
    #which counts as a digit in a higher base
    
    # if char is below 'A' or above 'Z', it is outside the range of these and is not capital letter
    li t5, 0x41
    blt t0, t5, not_upper
    li t5, 0x5A
    blt t5, t0, not_upper
    
    is_upper:
    li t5, 0x41
    sub t0, t0, t5
    addi t0, t0, 10
    j process_char
    
    not_upper:
    
    # if char is below 'a' or above 'z', it is outside the range of these and is not lowercase letter
    li t5, 0x61
    blt t0, t5, not_lower
    li t5, 0x7A
    blt t5, t0, not_lower
    
    is_lower:
    li t5, 0x61
    sub t0, t0, t5
    addi t0, t0, 10
    j process_char
    
    not_lower:
    
    # if we have reached this point, result invalid and end function
    # this is only reached if the byte was not a valid digit or alphabet character
    j strint_end_error
    
    process_char:
    
    blt t2, t0 strint_end_error #if this value is above or equal to radix, it is too high despite being a valid digit/alpha
    
    mul s0, s0, t2 # multiply s0 by the radix
    add s0, s0, t0 # add the correct value of this digit
    
    j read_strint # jump back and continue the loop if nothing has exited it
    
    strint_end_error:  #we jump here if there was an error with one of the chars
    addi a0, a0, 1 #add 1 to the a0 register indicating an error occurred
    
    strint_end: #we jump here when no errors happened
    ret
    
    ###############################################################################
    # This putstr function is my most portable function for RISC-V simulators     #
    # It calculates the length of a zero terminated string before printing it     #
    # This is the same way used in my Intel Assembly programs for DOS and Linux   #
    # This function was written to operate the same in both RARS and riscemu      #
    ###############################################################################
    
    putstr:
    
    mv t1, s0                       # t1 will be used as an index register
    
    putstr_strlen_start:
    lb t0, 0(t1)                    # load byte into t0 from address of t1
    beq t0, zero, putstr_strlen_end # if t0==0, then we jump to the end of the loop.
    addi t1, t1, 1                  # go to next byte
    j putstr_strlen_start           # jump to start of the loop
    putstr_strlen_end:              
    
    li a0, 1                        # STDOUT file number
    mv a1, s0                       # address of string 
    sub a2, t1, s0                  # length of string
    li a7, 64                       # write call number
    ecall                           # environment call
    
    ret
    
    #############################################################################
    # The next four 3 functions print things to standard output                 #
    # All of them use the putstr function above to achieve the output           #
    # They use the stack to preserve the values of the s0 and t1 registers used #
    # They also use global variables in the data section                        #
    #############################################################################
    
    #the putchar function, which is named after the C language function of the same name
    #prints the lowest byte of the s0 register as a byte or character to standard output
    
    putchar:
    
    addi sp, sp, -12
    sw ra, 0(sp)
    sw s0, 4(sp)
    sw t1, 8(sp)
    
    la t1, char
    sb s0, 0(t1)
    la s0, char
    jal putstr
    
    lw ra, 0(sp)
    lw s0, 4(sp)
    lw t1, 8(sp)
    addi sp, sp, 12
    
    ret
    
    # the putspace function prints a space to standard output
    
    putspace:
    
    addi sp, sp, -8
    sw ra, 0(sp)
    sw s0, 4(sp)
    
    la s0, space
    jal putstr
    
    lw ra, 0(sp)
    lw s0, 4(sp)
    addi sp, sp, 8
    
    ret
    
    # the putline function prints a newline to standard output
    
    putline:
    
    addi sp, sp, -8
    sw ra, 0(sp)
    sw s0, 4(sp)
    
    la s0, line
    jal putstr
    
    lw ra, 0(sp)
    lw s0, 4(sp)
    addi sp, sp, 8
    
    ret
    
    ##########################################################################
    # chastdin extension functions                                           #
    #                                                                        #
    # all functions that deal with getting strings and characters from stdin #
    ##########################################################################
    
    # the getstr function will read a string into a buffer from stdin
    # and return it in the s0 register for printing with the putstr function
    # the (count) variable will also return the number of characters
    
    getstr:
    
    li t0, 0                        # use t0 register to track chars read
    la a1, buf                      # load address of buffer for read string
    li a2, 1                        # read only 1 byte for each env call
    
    getstr_chars:
    
    li a0, 0                        # STDIN file number
    li a7, 63                       # read call number
    ecall                           # environment call
    
    # Branch to label getstr_end if a0 is less than a2
    # a0 is the return value of this environment read call
    # as will be -1 on error or 1 if successful
    # because we read 1 character at a time
    
    blt a0, a2, getstr_end
    
    # if no error, test range of the last byte
    
    lb t1, 0(a1)      #load byte at address (a1) into t1 register
    
    # if t1 is less than 0x21
    # or t1 is more than 0x7E
    # branch to function end because it is outside of print range
    
    li t2, 0x21
    blt t1, t2, getstr_end
    li t2, 0x7E
    blt t2, t1, getstr_end
    
    # otherwise, proceed to read more characters
    add t0, t0, a0    # add to read counter
    addi a1, a1, 1    # add 1 to buffer pointer register a1
    j getstr_chars # unconditional jump to getstr_chars
    
    getstr_end:
    
    la t2, count       #load address of count into t2
    sw t0, 0(t2)       #store number of chars read at (count) address
    la t2, last_char   #load address of last_char into t2
    sb t1, 0(t2)       #store last byte at (last_char) address
    sb zero, 0(a1)     #store byte zero to terminate string
    la s0, buf         #return address of buf in s0 register
    
    ret
    
    
    
    
    # the getline function will read a string into a buffer from stdin
    # and return it in the s0 register for printing with the putstr function
    # the (count) variable will also return the number of characters
    # this function will get the whole line including spaces
    
    getline:
    
    li t0, 0                        # use t0 register to track chars read
    la a1, buf                      # load address of buffer for read string
    li a2, 1                        # read only 1 byte for each env call
    
    getline_chars:
    
    li a0, 0                        # STDIN file number
    li a7, 63                       # read call number
    ecall                           # environment call
    
    # Branch to label getline_end if a0 is less than a2
    # a0 is the return value of this environment read call
    # as will be -1 on error or 1 if successful
    # because we read 1 character at a time
    
    blt a0, a2, getline_end
    
    # if no error, test range of the last byte
    
    lb t1, 0(a1)      #load byte at address (a1) into t1 register
    
    # if t1 is less than 0x20
    # or t1 is more than 0x7E
    # branch to function end because it is outside of print range
    
    li t2, 0x20
    blt t1, t2, getline_end
    li t2, 0x7E
    blt t2, t1, getline_end
    
    # otherwise, proceed to read more characters
    add t0, t0, a0    # add to read counter
    addi a1, a1, 1    # add 1 to buffer pointer register a1
    j getline_chars # unconditional jump to getline_chars
    
    getline_end:
    
    la t2, count       #load address of count into t2
    sw t0, 0(t2)       #store number of chars read at (count) address
    la t2, last_char   #load address of last_char into t2
    sb t1, 0(t2)       #store last byte at (last_char) address
    sb zero, 0(a1)     #store byte zero to terminate string
    la s0, buf         #return address of buf in s0 register
    
    ret
    
    
    
    # Short Description of strlen:
    # The strlen function gets the length of string in s0 and returns it in s0
    # This is the same algorithm used in my putstr function but is independent of an operating system.
    
    strlen:
    
    mv t1, s0                       # t1 will be used as an index register
    
    strlen_start:
    lb t0, 0(t1)                    # load byte into t0 from address of t1
    beq t0, zero, strlen_end        # if t0==0, then we jump to the end of the loop.
    addi t1, t1, 1                  # go to next byte
    j strlen_start                  # jump to start of the loop
    strlen_end:              
    
    sub s0, t1, s0                  # return length of string in s0
    
    ret
    
    
    # Short Description of strcmp:
    # strcmp compares the string at s0 to the one at s1
    # t0 returns 0 if the strings are the same and non zero if different
    # the algorithm is simple but I will explain it for those who are confused
    
    # Long Description of strcmp:
    # each byte from each string is loaded into the t0 and t1 registers
    # the bytes are compared. if they are different, then we jump to the end
    # However, if they are the same, then we check if one of them is zero
    # if it is zero, this also jumps to the end of the function
    # If neither jump took place, then we jump to the start of the loop
    # but when the function finally ends t1 will be subtracted from t0
    # this ensures that the t0 register returns zero if the final characters are the same
    # a zero result in t0 also guarantees that both strings are equal
    
    strcmp:
    
    mv a0, s0 # move pointer s0 to a0
    mv a1, s1 # move pointer s1 to a1
    
    strcmp_start:
    
    #read a byte from each string
    lb t0, 0(a0) 
    lb t1, 0(a1) 
    #if the two bytes are not equal end comparison
    bne t0, t1, strcmp_end
    
    #but if they are equal, test for zero
    #if one of them is zero, also end the loop
    beq t0, zero, strcmp_end
    
    addi a0, a0, 1                  # go to next byte
    addi a1, a1, 1                  # go to next byte
    
    j strcmp_start
    
    strcmp_end:
    
    #subtract t1 from t0
    #if t0 is still zero after the function returns
    #it means that the strings are equal
    sub t0, t0, t1
    
    ret
    
    
  • chastdin for FreeBASIC

    I have rewritten my chastdin program (the stack based calculator that reads keyboard input from standard input) into the FreeBASIC programming language. I did it as an exercise to prepare myself for a future book on the BASIC programming language which was my first programming language. FreeBASIC is compatible with QBASIC which is what I first started on. Luckily, BASIC is not so different from C but I had to spend a lot of time on the documentation to refresh my memory in how I used to do things with it.

    Eventually I would like to make a GUI version of this command line calculator. I am trying to take baby steps in working my way into programs that the average person would use. However, command line utilities are still the easiest to build and that is an acceptable place to start.

    main.bas

    #include "chastelib.bi"
    #include "chastdin.bi"
    
    dim shared as integer chastack(256)
    dim shared as integer csi=0 'Chastity's Stack Index
    
    radix=10
    
    dim as integer a,b
    dim shared as string s
    
    sub help()
    ?  "chastdin is a stack based interactive calculator"
    ?  "Numbers are pushed on the stack and commands can do math."
    ?  "It is a fork of chastack that reads from stdin instead of arguments."
    ?  "Each line can contain multiple numbers or commands."
    ?
    ?  "Math commands are add,sub,mul,div,rem"
    ?  "And use the top two stack numbers for their operations"
    ?
    ?  "The setradix command uses the top of stack as the new radix"
    ?  "The exit command ends the program"
    ?  "The ? command prints the entire stack"
    ?
    end sub
    
    sub stack_check()
     if csi>0 then
      chastack(csi+1)=0 /'erase old top of stack because command was successful'/
     else
      print "Error: two numbers required for command: ";s
      csi+=1 /'increment the pointer to what it was before the failed command'/
     end if
    end sub
    
    help():
    
    while s<>"exit"
    
    s=""
    
     s=getstr() 'read and ignore empty strings
    
    'print entire stack
    if s="?" or s="print" then
     b=csi
     while csi>0
      print intstr(chastack(csi))
      csi-=1
     wend
     csi=b
    
    elseif s="exit" then
    exit while
    
    elseif s="help" then
    help()
    
    elseif s="setradix" then
     if csi>0 then
     radix=chastack(csi)
     chastack(csi)=0
     csi-=1
     else
      print "Error: need one number on stack for command: ";s
     end if
    
    elseif s="add" then
    b=chastack(csi)
    csi-=1
    a=chastack(csi)
    a+=b
    chastack(csi)=a
    stack_check()
    
    elseif s="sub" then
    b=chastack(csi)
    csi-=1
    a=chastack(csi)
    a-=b
    chastack(csi)=a
    stack_check()
    
    elseif s="mul" then
    b=chastack(csi)
    csi-=1
    a=chastack(csi)
    a*=b
    chastack(csi)=a
    stack_check()
    
    elseif s="div" then
    b=chastack(csi)
    csi-=1
    a=chastack(csi)
    a\=b
    chastack(csi)=a
    stack_check()
    
    elseif s="rem" then
    b=chastack(csi)
    csi-=1
    a=chastack(csi)
    a=a mod b
    chastack(csi)=a
    stack_check()
    
    else
    
    'try to interpret string as a number if not empty
     a=strint(s)
     if strint_errors<>0 or len(s)=0 then
     'print s;" cannot be added to the stack because it is not a valid number"
     else
     csi+=1
     chastack(csi)=a
     print intstr(a);" was added to the stack"
     end if
    
    end if
    
    wend
    
    /'
     This is a FreeBASIC program.
    
     compile and run as:
    
     fbc main.bas && ./main
    '/
    
    

    chastelib.bi

    /'
     global variables to define radix and formatting
     for the intstr function
    '/
    dim shared as integer radix=2
    dim shared as integer int_width=1
    
    /'
     translation of intstr function for FreeBASIC
     by original C programmer Chastity White Rose
    '/
    function intstr(i as uinteger) as string
     dim as string s=""
     dim as integer w=0
     dim as byte c
    
     while i<>0 or w<int_width 
    
      c=i mod radix                  
      i\=radix                     
    
      if c<10 then 
      c+=48
      else
      c+=55
      end if
    
      s=chr(c)+s
    
      w+=1                     
     wend
    
    return s
    end function
    
    /'
     global variable for error detection in strint function
     this variable will be zero if last string was a number
    '/
    dim shared as integer strint_errors=0
    
    /'
     translation of strint function for FreeBASIC
     by original C programmer Chastity White Rose
    '/
    function strint(s as string) as uinteger
    dim as uinteger i=0
    dim as integer x=0,y=len(s)
    dim as byte c
    
    strint_errors = 0 /' clear errors '/
    
    while x<y
    
     /' read digit from string '/
     c=s[x]
    
     /' 0 to 9 '/
     if c >= 48 and c <= 57 then
     c-=48
     /' A to Z '/
     elseif c >= 65 and c <= 90 then
     c-=65
     c+=10
     /' a to z '/
     elseif c >= 97 and c <= 122 then
     c-=97
     c+=10
     /' whitespace '/
     elseif c >= 0 and c <= 32 then
      exit while /' exit correctly at string end '/
     else
      strint_errors+=1
      print "Error: ";chr(s[x]);" is not an alphanumeric character!"
      exit while /' exit at invalid character '/
     end if
    
     if c>=radix then
      strint_errors+=1
      print "Error: ";chr(s[x]);" is not a valid character for radix ";radix
      exit while /' exit at digit wrong for radix '/
     end if
    
     /'multiply by radix then add digit'/
     i*=radix
     i+=c
    
    x+=1
    wend
    
    return i
    end function
    

    chastdin.bi

    dim shared as string stdin_buf
    dim shared as integer stdin_buf_index
    dim shared as integer stdin_buf_length=0
    
    function getstr() as string
    dim as string s=""         'create empty string
    dim as byte c              'temporary byte/char variable
    
    /'
    this section gets a line of text
    if the length of the string/buffer is 0
    '/
    
    if stdin_buf_length=0 then      'check if there are characters in the buf
    input "-> ",stdin_buf           'if not, read a line of text
    stdin_buf_index=0               'set index to zero
    stdin_buf_length=len(stdin_buf) 'set the length
    end if
    
    /'
    regardless of whether input was added above
    or if it still had bytes from the last input
    we then extract characters one at a time into the
    substring s to be returned from the function
    '/
    
    while stdin_buf_index<stdin_buf_length
    c=stdin_buf[stdin_buf_index]
    stdin_buf_index+=1
    if(c>=33) and (c<=126) then
    s=s+chr(c)
    else
    exit while
    endif
    wend
    
    /'
    if the index matches the length of buffer
    set length to zero so that more will be read
    next time this function is called
    '/
    
    if stdin_buf_index=stdin_buf_length then
    stdin_buf_length=0
    end if
    
    return s
    end function
    
    /'
    the getline function always gets an entire line of text
    I don't really need it but it is here as a reminder of
    how to use the input statement in FreeBASIC
    '/
    
    function getline() as string
    dim as string s=""
    input "-> ",stdin_buf
    s=stdin_buf
    return s
    end function
    

  • chastelib for Pascal Programming Language

    I managed to hack my four functions from chastelib into the Pascal programming language. This program includes the functions and the test suite which works just like the C version. The code is a bit more complex than the C version because strings and characters are handled very differently than they are in the C programming language. The strint function was the hardest to write but it seems to be working according to the standards I require.

    I am doing this for education and possibly a future book on old programming languages. Pascal is nice but I will also be studying BASIC again.

    program chastelib;
    
    const
     string0='Official test suite for the Pascal version of chastelib.'#10;
    
    var //this is the global variable section
     a:integer;
     b:integer;
     
     radix:integer;       //current radix being used
     int_width:integer=1; //global integer width
     strint_errors:integer=0; //error result for strint function
    
    (*
    A function to print a string using Pascal's write function.
    *)
    procedure putstr(s:string);
    begin
     write(s);
    end;
    
    (*
     a function to return a string form of an integer
     using the global radix variable
    *)
    function intstr(i:integer):string;
    var
     s:string=''; //string that will be built and returned from this function
     width:integer=0; //the current width
     c:integer;
     ch:char;
    begin
     while (i>0) or (width<int_width) do
     begin
    
      c:=i mod radix; //get integer division modulus or remainder
      i:=i div radix; //get integer division quotient
    
      (*turn remainder c into character ch for digit in this radix*)
      if c<10 then
      begin
       ch:=chr(c+48);
      end
      else
      begin
       ch:=chr(c+55);
      end;
    
       s:=ch+s; //prefix the string with this character
       width+=1;
    
     end;
    
     intstr:=s; //return this string from the function
    
    end;
    
    (*use both putstr and intstr to print an integer*)
    procedure putint(i:integer);
    begin
     putstr(intstr(i));
    end;
    
    (*
    Because characters and integers are separate types in Pascal,
    it is required to get the ASCII value of characters in the string
    for the strint function so I can do the math the same way
    as I did in the C version of the function.
    *)
    
    function strint(s:string):integer;
    var
     i:integer=0; //integer that will be built and returned from this function
     x:integer=1; //index used to scan forward through the string
     c:integer=0;
    begin
     strint_errors := 0; (*set zero errors before we parse the string*)
     if (radix<2) or (radix>36 ) then
     begin
      strint_errors+=1;
      writeln('Error: radix ',radix,' is out of range!');
     end;
     while(x<=length(s)) do
     begin
      c:=ord(s[x]);
      if (c>=ord('0')) and (c<=ord('9')) then 
      begin
       c-=ord('0')
      end
      else if (c>=ord('A')) and (c<=ord('Z')) then
      begin
       c-=ord('A');
       c+=10;
      end
      else if (c>=ord('a')) and (c<=ord('z')) then
      begin
       c-=ord('a');
       c+=10;
      end
      
      else if (c < $21 ) then
      begin
       break; (*end loop because we have found whitespace*)
      end
      
      else
      begin
       strint_errors+=1;
       writeln('Error: ',s[x],' is not an alphanumeric character!');break;
      end;
      
      if(c>=radix) then
      begin
       strint_errors+=1;
       writeln('Error: ',s[x],' is not a valid character for radix ',radix);
       break;
      end;
      
      i*=radix; //multiply by the radix
      i+=c;     //add the digit from the character processed
    
      x:=x+1;
     end;
     strint:=i;
    end;
    
    
    
    
    begin
     radix:=16; //set the radix used by both intstr and strint functions
    
     a:=0;
     b:=strint('100');
    
     putstr(string0);
    
     while a<b do
     begin
      radix:=2;
      int_width:=8;
      putint(a);
      putstr(' ');
      radix:=16;
      int_width:=2;
      putint(a);
      putstr(' ');
      radix:=10;
      int_width:=3;
      putint(a);
    
      if (a>=$20) and (a<=$7E) then
      begin
       putstr(' ');
       putstr(chr(a));
      end;
    
      putstr(#10);
      a+=1;
     end;
     
     putstr(string0);
    
    end.
    
    (*
     fpc main.pas && ./main
    *)
    
  • new program: chastdin

    I wrote another program which is actually a modification of chastack. This gets input from a user while it is running. Despite how simple it may seem, I had to work at reading from the keyboard because there are multiple ways to read a string from the user. I may add more to this program later, but it has all the important functions of a stack based calculator. Here is a screenshot that shows me using it. You can probably figure out what the commands do based on their name and the numbers printed. I have also attached the full assembly source code to this post.

    main.asm

    format ELF executable
    entry main
    
    include 'chastelib32.asm'
    include 'chastdin32.asm'
    
    main:
    
    mov dword[radix],10    ;I can choose the radix for integer output!
    mov dword[int_width],1 ;and the width of each integer for padded zeros
    
    mov ebp,chastack       ;mov the address of the beginning of the stack to ebp registers
    
    ;this program does not read command line arguments
    ;it always displays a message to tell user what the program does
    mov eax,string_help
    call putstring
    
    mov [last_char],0xA ;set newline as last_char so prompt will display
    
    main_loop:
    
    ;show the arrow indicating we wait for the user to enter something
    ;but only show it when the last character is a newline
    ;otherwise it will print too many if multiple commands were entered on the same line
    cmp [last_char],0xA
    jnz skip_prompt
    mov eax,string_prompt
    call putstring
    skip_prompt:
    
    call getstring ;get string and return address in eax
    
    ;we must restart the loop in case of an empty string
    ;if we didn't, strint would read the empty string and return 0
    ;then zero would be pushed to the stack, which is not what we want
    
    cmp dword[count],0 ;were there zero characters read?
    jz main_loop ;if yes, this was an empty string, retry input
    
    mov esi,eax    ;mov string to esi for string comparison
    
    ;Now we process the string the user entered
    ;First, we will try testing for commands
    ;If any of the predefined strings match the string in esi
    ;We jump to the label for that command
    
    mov edi,string_add
    call strcmp
    jz command_add
    
    mov edi,string_sub
    call strcmp
    jz command_sub
    
    mov edi,string_mul
    call strcmp
    jz command_mul
    
    mov edi,string_div
    call strcmp
    jz command_div
    
    mov edi,string_rem
    call strcmp
    jz command_rem
    
    mov edi,string_query
    call strcmp
    jz command_query
    
    mov edi,string_clear
    call strcmp
    jz command_clear
    
    mov edi,string_exit
    call strcmp
    jz command_exit
    
    ;The default command is to turn the argument into a number and push to stack
    command_num:
    
    mov eax,esi          ;mov the string to eax for processing numbers
    call strint          ;try to get a number from the string pointed to by eax
    cmp [strint_error],0 ;did we have zero errors in the strint function?
    jz num_push          ;if there were no errors, push this to stack
    
    mov eax,string_err
    call putstring
    mov eax,esi
    call putstring
    call putline
    jmp num_push_end ;skip the push because this can't be used
    
    num_push:        ;push the number to the fake stack
    add ebp,4
    mov [ebp],eax
    num_push_end:
    jmp main_loop
    
    ;These are the labels and code for each of the commands
    ;When a command is done, we jump back to the beginning of the loop
    
    command_add:
    mov eax,[ebp]
    mov dword[ebp],0
    sub ebp,4
    add [ebp],eax
    jmp main_loop
    
    command_sub:
    mov eax,[ebp]
    mov dword[ebp],0
    sub ebp,4
    sub [ebp],eax
    jmp main_loop
    
    command_mul:
    mov ebx,[ebp]
    mov dword[ebp],0
    sub ebp,4
    mov eax,[ebp]
    mov edx,0     ;zero edx before multiply
    mul ebx       ;multiply eax with value in ebx
    mov [ebp],eax
    jmp main_loop
    
    command_div:
    mov ebx,[ebp]
    mov dword[ebp],0
    sub ebp,4
    mov eax,[ebp]
    mov edx,0 ;zero edx before divide
    div ebx   ;divide eax with value in ebx
    mov [ebp],eax ;store quotient on stack
    jmp main_loop
    
    command_rem:
    mov ebx,[ebp]
    mov dword[ebp],0
    sub ebp,4
    mov eax,[ebp]
    mov edx,0 ;zero edx before divide
    div ebx   ;divide eax with value in ebx
    mov [ebp],edx ;store remainder on stack
    jmp main_loop
    
    command_query: ;print all numbers on the stack
    push ebp ;save value of ebp
    command_query_loop:
    cmp ebp,chastack ;is ebp equal to the address of stack start?
    jz command_query_end  ;if it is, end the putstack loop
    mov eax,[ebp]
    sub ebp,4
    call putint_and_line
    jmp command_query_loop
    command_query_end:
    pop ebp ;restore ebp to what it was before this command
    jmp main_loop
    
    command_clear: ;erase all numbers on the stack
    command_clear_loop:
    cmp ebp,chastack ;is ebp equal to the address of stack start?
    jz command_clear_end  ;if it is, end the putstack loop
    mov dword[ebp],0
    sub ebp,4
    jmp command_clear_loop
    command_clear_end:
    jmp main_loop
    
    command_exit: ;end the program
    
    main_loop_end:
    
    mov eax,1        ;exit (kernel opcode 1 on 32 bit systems)
    mov ebx,0        ;return 0 status on exit - 'No Errors'
    int 80h          ;system call for 32-bit Linux kernel
    
    argc dd 0
    
    string_err db 'Error: invalid number or command: ',0 ;Generic error message
    string_add db 'add',0
    string_sub db 'sub',0
    string_mul db 'mul',0
    string_div db 'div',0
    string_rem db 'rem',0
    string_exit db 'exit',0
    string_query db '?',0
    string_clear db 'clear',0
    
    string_prompt db '-> ',0
    
    string_help db 'chastdin is a stack based interactive calculator',0xA
                db 'Numbers are pushed on the stack and commands can do math.',0xA
                db 'It is a fork of chastack that reads from stdin instead of arguments.',0xA
                db 'Each line can contain multiple numbers or commands.',0xA
                db 'Math commands are add,sub,mul,div,rem',0xA
                db 'The exit command ends the program',0xA
                db 'The ? command prints the entire stack',0xA,0xA,0
    
    ;This program uses a virtual stack for convenience and portability
    ;I allocate memory for a virtual stack that we can index as if it was the real stack
    ;I name it "chastack" for Chastity's stack.
    
    db 6 dup 0 ;extra padding bytes
    chastack: rd 0x100
    

    chastdin32.asm

    ;Chastity's Standard Input header file
    ;The functions here are designed to read strings and numbers from standard input.
    
    ;getstring ;read characters from stdin until the first whitespace
    ;getline   ;read characters from stdin until the first newline,EOF,tab,etc.
    ;strcmp    ;compare two strings similar to the same function in C
    
    ;these variables are used as the default controllers
    ;for the getstring and getline functions
    ;buf stores keyboard input during those functions
    ;count stores how many bytes were read
    ;last_char stores the last character read
    ;usually this will be a space, tab, or newline
    
    buf db 0x100 dup '?'
    count dd 0
    last_char db 0
    
    ;summary
    ;the getstring function is the reverse function of putstring
    ;instead of printing a string to standard output
    ;it reads a string from standard input (AKA the keyboard)
    
    ;details
    ;the getstring function is designed to get a string of text
    ;which is terminated by whitespace or any non printable character
    ;the idea is that multiple strings can be passed on one line
    ;separated by spaces, similar to command line arguments
    ;this function was written for the specific purpose of converting any of
    ;my programs that used command line arguments to read from stdin instead
    
    getstring:
    
    mov [count],0 ;set count of characters read during this function to zero
    mov edx,1     ;number of bytes to read
    mov ecx,buf   ;address to store the bytes
    
    getstring_chars:
    
    mov ebx,0     ;read from stdin
    mov eax,3     ;invoke SYS_READ (kernel opcode 3)
    int 80h       ;call the kernel
    
    cmp eax,1     ;was 1 character read?
    jnz getstring_end ; if not, then end this loop
    
    mov al,[ecx]  ;mov last character read into al register
    
    ;check if this character is in the proper range to be part of the string
    
    cmp al,0x21      ;compare with 0x21 (!=exclamation)
    jb getstring_end ;jump if below to getstring_end label
    cmp al,0x7E      ;compare with 0x7E (tilde)
    ja getstring_end ;jump if above to getstring_end label
    
    ;if neither jump happened, keep the character and
    
    inc [count]   ;increment how many characters we have read
    inc ecx       ;increment address where next byte is read from
    jmp getstring_chars ;jump back to start of loop and keep reading
    
    getstring_end:
    
    mov [last_char],al ;save the last character read
    mov byte[ecx],0 ;terminate this string with a zero
    
    mov eax,buf ;mov the buffer address to eax for returning the string
    
    ret
    
    ;the getline function gets an entire line of text from the keyboard
    ;calling this function allows for a string that can contain spaces
    ;it considers as anything outside the range of 0x20 to 0x7E as the end of line character
    ;this is because the end of the line might be 0x0A on Linux
    ;or it might be 0x0D,0x0A on DOS or Windows.
    ;technically, it means tab will also terminate a line
    ;the intended use of this function is to read a filename
    ;filenames can contain spaces
    
    getline:
    
    mov [count],0 ;set count of characters read during this function to zero
    mov edx,1     ;number of bytes to read
    mov ecx,buf   ;address to store the bytes
    
    getline_chars:
    
    mov ebx,0     ;read from stdin
    mov eax,3     ;invoke SYS_READ (kernel opcode 3)
    int 80h       ;call the kernel
    
    cmp eax,1     ;was 1 character read?
    jnz getline_end ; if not, then end this loop
    
    mov al,[ecx]  ;mov last character read into al register
    
    ;check if this character is in the proper range to be part of the string
    
    cmp al,0x20    ;compare with 0x20 (space)
    jb getline_end ;jump if below to getstring_end label
    cmp al,0x7E    ;compare with 0x7E (tilde)
    ja getline_end ;jump if above to getstring_end label
    
    ;if neither jump happened, keep the character and
    
    inc [count]       ;increment how many characters we have read
    inc ecx           ;increment address where next byte is read from
    jmp getline_chars ;jump back to start of loop and keep reading
    
    getline_end:
    
    mov byte[ecx],0 ;terminate this string with a zero
    
    mov eax,buf ;mov the buffer address to eax for returning the string
    
    ret
    
    ;summary
    ;strcmp compares the string at esi to the one at edi
    ;eax returns 0 if the strings are the same and 1 if different
    ;the algorithm is simple but I will explain it for those who are confused
    
    ;details
    ;eax is initialized to zero
    ;a byte from each string is loaded into the al and bl registers
    ;the bytes are compared. if they are different, then we jump to the end
    ;However, if they are the same, then we check if one of them is zero
    ;for this purpose it doesn't matter whether we compare al or bl with zero
    ;because it is known that they are the same if the jnz did not take place
    ;if it is zero, this also jumps to the end of the function
    ;If neither jump took place, then we jump to the start of the loop
    ;but when the function finally ends bl will be subtracted from al
    ;this ensures that the function returns zero if the final characters are the same
    ;ebx,esi,and edi are preserved but eax is the return value
    ;also, the sub instruction at the end of the function also updates the flags
    ;so you can "jz" or "jnz" to a label after calling this function based on results
    
    strcmp:
    
    push ebx
    push esi
    push edi
    
    mov eax,0
    
    strcmp_start:
    
    ;read a byte from each string
    mov al,[edi]
    mov bl,[esi]
    cmp al,bl
    jnz strcmp_end
    
    cmp al,0
    jz strcmp_end
    
    inc edi
    inc esi
    
    jmp strcmp_start
    
    strcmp_end:
    sub al,bl
    
    pop edi
    pop esi
    pop ebx
    
    ret
    
  • AAA Linux: Chapter 15: chastecmp

    This post is a chapter from my recently published Linux edition of Assembly Arithmetic Algorithms. The story behind why I wrote the program featured in chapter 15 goes back to when I discovered how to cheat at video games. This story is worth sharing to inspire the next generation of gamers to learn computer math and programming just as I did.

    Chapter 15: chastecmp

    In this chapter, I will show you the source code of a file comparison program. This program is meant to find which bytes are different between two files that are similar but contain a few differences.

    I will use text files for my examples in this chapter, but the program actually does a binary file comparison and displays the different bytes in hexadecimal because it is a universally understood shorthand for binary that most C and Assembly programmers are already familiar with.

    First, here is the source code of chastecmp, which is the short name for “Chastity’s Comparison tool”. The name is also meant to refer to the “cmp” instruction, which is used a lot more in this program because it is essential.

    FASM chastecmp source

    ;Linux 32-bit Assembly Source for chastecmp
    format ELF executable
    
    main:
    
    ;radix will be 16 because this whole program is about hexadecimal
    mov dword[radix],16 ; can choose radix for integer input/output!
    mov dword[int_width],1
    
    pop eax ;get the number of arguments
    dec eax ;subtract 1 because we will ignore the name of the program
    pop ebx ;pop program name into a register to delete it from stack
    
    cmp eax,2 ;do we have two arguments to be used as filenames?
    jb help
    mov dword[offset],0 ;assume the offset is 0,beginning of file
    jmp arg_open_file_1
    
    help:
    mov eax,help_message
    call putstring
    jmp main_end
    
    arg_open_file_1:
    pop eax
    mov [filename1],eax ; save the name of the file we will open to read
    call putstring ;print the name of the file we will try opening
    
    mov ecx,0   ;open file in read mode 
    mov ebx,eax ;move filename for system call
    mov eax,5   ;invoke SYS_OPEN (kernel opcode 5)
    int 80h     ;call the kernel
    
    cmp eax,0
    js file_error_display ;end program if the file can't be opened
    mov [fd1],eax ; save the file descriptor number for later use
    mov eax,file_open
    call putstr_and_line
    
    arg_open_file_2:
    pop eax
    mov [filename2],eax ; save the name of the file we will open to read
    
    call putstring ;print the name of the file we will try opening
    
    mov ecx,0   ;open file in read mode 
    mov ebx,eax ;move filename for system call
    mov eax,5   ;invoke SYS_OPEN (kernel opcode 5)
    int 80h     ;call the kernel
    
    cmp eax,0
    js file_error_display ;end program if the file can't be opened
    mov [fd2],eax ; save the file descriptor number for later use
    mov eax,file_open
    call putstr_and_line
    
    files_compare:
    
    file_1_read_one_byte:
    mov edx,1       ;number of bytes to read
    mov ecx,buf1    ;address to store the bytes
    mov ebx,[fd1]   ;move the opened file descriptor into EBX
    mov eax,3       ;invoke SYS_READ (kernel opcode 3)
    int 80h         ;call the kernel
    
    ;eax will have the number of byte read after system call
    mov [count1],eax ;we save the number of byte read for later
    cmp eax,0
    jnz file_2_read_one_byte ;unless zero bytes were read, proceed to read from next file
    
    mov eax,[filename1]
    call putstring
    mov eax,end_of_file_string
    call putstr_and_line
    
    ;Even if we have reached the end of the first file,
    ;we still proceed to read a byte from the second file
    ;to see if it also ends at the same address
    
    file_2_read_one_byte:
    mov edx,1       ;number of byte to read
    mov ecx,buf2    ;address to store the bytes
    mov ebx,[fd2]   ;move the opened file descriptor into EBX
    mov eax,3       ;invoke SYS_READ (kernel opcode 3)
    int 80h         ;call the kernel
    
    ;eax will have the number of bytes read after system call
    mov [count2],eax ;we save the number of bytes read for later
    cmp eax,0
    jnz check_both_bytes ;unless zero bytes were read, proceed to compare bytes from both files
    
    mov eax,[filename2]
    call putstring
    mov eax,end_of_file_string
    call putstr_and_line
    
    jmp main_end ;we have reach end of one file and should end program
    
    check_both_bytes:
    
    ;we add the number of bytes read from both files
    mov eax,[count1]
    add eax,[count2]
    cmp eax,2
    jnz main_end
    
    compare_bytes:
    
    mov al,[buf1]
    mov bl,[buf2]
    
    ;compare the two bytes and skip printing them if they are the same
    cmp al,bl
    jz bytes_are_same
    
    ;print the address and the bytes at that address
    mov eax,[offset]
    mov dword[int_width],8
    call putint_and_space
    mov dword[int_width],2
    mov eax,0
    mov al,[buf1]
    call putint_and_space
    mov al,[buf2]
    call putint_and_line
    
    bytes_are_same:
    
    inc dword[offset]
    
    jmp files_compare
    
    file_error_display:
    
    mov eax,file_error
    call putstr_and_line
    
    main_end:
    
    ;this is the end of the program
    ;we close the open files and then use the exit call
    
    mov ebx,[fd1] ;file number to close
    mov eax,6   ;invoke SYS_CLOSE (kernel opcode 6)
    int 80h     ;call the kernel
    
    mov ebx,[fd2] ;file number to close
    mov eax,6   ;invoke SYS_CLOSE (kernel opcode 6)
    int 80h     ;call the kernel
    
    mov eax, 1  ; invoke SYS_EXIT (kernel opcode 1)
    mov ebx, 0  ; return 0 status on exit - 'No Errors'
    int 80h
    
    include 'chastelib32.asm'
    
    ;variables for displaying information
    help_message db 'chastecmp by Chastity White Rose',0Ah,0Ah
    db 9,'chastecmp file1 file2',0Ah,0Ah
    db 'Differing bytes are shown in hexadecimal',0Ah
    db 'until the EOF has been reached.',0Ah,0
    
    file_open db ' opened',0
    file_error db ' error',0
    end_of_file_string db ' EOF',0
    
    db 23 dup 0 ;fill with extra space to match 1024 executable size
    
    ;variables for managing files
    filename1 dd ? ;name of the file to be opened
    filename2 dd ? ;name of the file to be opened
    fd1 dd ?       ;file descriptor 1
    fd2 dd ?       ;file descriptor 2
    buf1 db ?      ;store byte from file 1 here
    buf2 db ?      ;store byte from file 2 here
    count1 dd ?    
    count2 dd ?
    offset dd ?
    

    How to use chastecmp

    Using the chastecmp program requires two filenames to be passed as command-line arguments. Although you can use any files you have, it makes sense to use a simple example with text files because they are so easy to create with the echo command.

    Run these commands to create the two files.

    echo "chandler is my birth name" > file1.txt
    echo "chastity is my trans name" > file2.txt
    

    Now that the files exist

    ./main file1.txt file2.txt
    

    If you have created these files and run the chastecmp program on them, you will see this result:

    file1.txt opened
    file2.txt opened
    00000003 6E 73
    00000004 64 74
    00000005 6C 69
    00000006 65 74
    00000007 72 79
    0000000F 62 74
    00000010 69 72
    00000011 72 61
    00000012 74 6E
    00000013 68 73
    file1.txt EOF
    file2.txt EOF
    

    How does chastecmp work?

    This program is much simpler than chastack or chastext, but it is close to 180 lines and still has some logic to follow. First thing it does is check to see how many command-line arguments were passed to the program. Since the name of the program always counts as 1, we subtract from this number and also pop the next argument into ebx just to get rid of it. The actual register used doesn’t matter in this case as long as it is not eax, which holds the number of arguments.

    The eax register is compared with 2. If this number is below 2, then there are not enough arguments to continue the program, and it will end. Otherwise, it will proceed to use the open call with both filenames and assume these files exist. If they do not exist, it will print the filename and then say error.

    If both files are opened, it will keep reading 1 byte from each file descriptor and store each in its own buffer of 1 byte. If the two bytes are the same, they will be ignored. However, if they are different, the address and the values of both bytes at that address will be displayed.

    The variable “offset” is used to keep track of which address we are at in both files, but it isn’t used to lseek in this program because we are going from beginning to end.

    If at any time the read system call returns 0, a message is displayed with the filename and EOF to tell the user that the end of that file has been reached.

    In the example I just used, both files are the same length of 26 bytes and will reach the end at the same time.

    But why should I care?

    The average person probably does not know why it matters to see the hexadecimal differences between two files. I know it seems silly, especially for small text files as I used in this chapter’s examples. However, I can give two examples of times I have used this information.

    The first example is relevant to Chapter 2, where I presented the header file “chaste-elf-32.nasm” which can be included to make a loadable program using the NASM assembler.

    I read the specification document for ELF files to describe what the fields were named and what the values meant. However, this informational alone was not enough for me to successfully create the custom ELF header. I had to create ELF executable files with FASM because it has this feature built in. By creating slightly different programs, I was able to compare the binary differences in the different source files fed to FASM. The chastecmp program was extremely helpful to me as I used it hundreds of times in reverse engineering the ELF format.

    One of my discoveries was that when the size of a program increased, either by adding more code or adding more data statements, there was a number in the header that also increased. As it turns out, the memory size of the file increased even when data reservation keywords (such as rb,rw,rd, and rq) were used, even when the size of the file itself didn’t.

    The specification could tell me a lot, but without the example ELF headers FASM was already creating, I would not have been able to create dynamic headers to match programs written in FASM. I probably spent 12 hours on that project, but at least I can assembler any of my programs with NASM if I make the necessary syntax changes.

    But perhaps a more fun example, and also the reason I got started with programming, was that I used a file comparison tool to cheat at a Norse mythology game years ago. The game was called Castle of the Winds, and it ran on Windows 3.1, 98, and even XP.

    One of the features of that specific game was that it let you save the game at any time. I remember that I had 5 mana points. I saved the first file and then cast the magic arrow spell to spend one point. I then saved a second file and ran the Windows “fc” command to compare the two files in binary mode.

    fc /b 1.cwg 2.cwg
    

    It told me the address of the byte that had changed from 5 to 4. I then opened this in a hex editor named XVI32 and changed this byte to different values.

    In time, I was able to not only change my mana points but also hit points and experience points to make myself invincible in that game.

    I didn’t really know much about hexadecimal at this point, but by trial and error, I accidentally started understanding it. It was this experience of cheating in a video game that led me to learn about binary and hexadecimal number systems originally.

    I had seen for the first time that an understanding of computer arithmetic could allow me to break the rules and do things in a video game that the developer could not predict or prevent me from doing. In those days, I learned to do the same with many video games and had many fun adventures.

    In modern times, developers have gotten smarter and have put measures in place to prevent this form of cheating. Most notably, more games are multiplayer and read data from a server that stores the game data, where no user can hack it.

    But you have to understand that back in the 90s, nearly every single player game could be hacked that stored its data locally and didn’t connect to the internet. I have had people criticize my habit of cheating in single-player games and say that it ruins the experience of the game.

    But what they don’t understand is that I didn’t care about the video game I was hacking, because Arithmetic had become my favorite game. My love of math was so great that I learned computer programming and had more fun writing programs in BASIC, C, and Assembly than I did playing video games in the first place.

    I can’t hack most modern games with these tricks, but I have found the art of computer programming, which is much more satisfying than any video game I have played in my life.

    In summary, the chastecmp program does the same thing as the “fc /b” command from DOS and Windows did. When I switched to Linux as my primary operating system, I wrote my own file comparison tool to always keep the fond memories of my childhood with me.