I think it's because the new_uninit_slice call Rust will trigger a panic? Or abort? With little-to-no chance for recovery? (I know little about Rust.)
If so, I can see why someone that someone coming from Rust might consider exit() to be the appropriate solution for C, even for library code which should never be in charge of deciding how a program should exit.
I think the essay could be improved by highlighting the different worldviews.
I'm also old enough that
// Allocate enough room for the string and its null terminator.
char* reversed = malloc(len + 1);
makes me nervous. Even for char -- I've never been on a system where sizeof(char) != 1 -- I want to see the sizeof included in the calculation, like:
so I don't have to think about sizeof(char) being special.
As long as I'm here, I'm a bit confused about the purpose of the "char* error_message" in the proposed Result. Why a char* vs a const char * or even better, an int with an error code? Who sees the message? Do we expect they know English, or will they be localized? Will the error message text be frozen forever, or might it change in the future?
> I've never been on a system where sizeof(char) != 1
And you never will, since sizeof(char) is guaranteed to always be 1.
I'm guessing you were thinking of CHAR_BIT != 8, but even then I'm not sure it would make a difference since malloc takes its argument size in bytes and a char more or less is a byte in C.
(Consider that char*s are also how you access the byte-level representation of objects in C. If chars were not the minimum addressable unit then that use wouldn't work)
(And CHAR_BIT is required to be 8 by POSIX, so even though the language and some exotic hardware will have it at non-8, it’s exceedingly rare at this point.)
So I don't believe your statement "sizeof(char) is guaranteed to always be 1" is correct.
> chars are also how you access the byte-level representation of objects in C
Where does the spec say that a char can be used to address any point in an object?
There's all sorts of oddities like tagged architectures which the C spec handles which I know essentially nothing about, but which break common expectations about how C works. I believe this is one of them.
I believe the following is undefined behavior in C, even though your compiler may let you do it, at least sometimes, and on modern desktop hardware:
int i = 12345;
char *s = ((char *)&i) + 1;
char c = *s;
I believe the following is the correct (or less incorrect) way to do it:
char tmp[sizeof(int)];
memcpy(tmp, &i, sizeof(int));
char c = tmp[1];
Yes, but in C standardese a "byte" is not necessarily the 8 bits that it's normally thought to be these days. From C89 Section 2.2.4.2 Numerical Limits [-1]:
> maximum number of bits for smallest object that is not a bit-field (byte) CHAR_BIT 8
i.e., CHAR_BIT is the number of bits in a byte. C23 has a similar definition, and further defines CHAR_WIDTH that is defined to expand to the same value as CHAR_BIT.
> So I don't believe your statement "sizeof(char) is guaranteed to always be 1" is correct.
From C89 section 3.3.3.4 The sizeof operator [0]:
> When applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1.
This wording remains basically identical through C23 [1].
> Where does the spec say that a char* can be used to address any point in an object?
From C89 section 3.3 Expressions:
> An object shall have its stored value accessed only by an lvalue that has one of the following types:
> <snip>
> * a character type.
This also remains the case up through C23.
I think you're thinking of the strict aliasing rule with your example, but character types are one of the exceptions to said rule so I think your example is actually fully defined. It'd be UB if you casted to an incompatible type like a float, I believe.
I .... okay, my mind is blown. It absolutely says that, and https://en.cppreference.com/cpp/language/types adds "this allows the extreme case in which bytes are sized 64 bits, all types (including char) are 64 bits wide, and sizeof returns 1 for every type."
So that scaling factor is placed into CHAR_BIT, and would be 64 on that DSP compiler.
> An object shall have its stored value accessed only by an lvalue that has one of the following types:
Yes, my confusion comes down to my confusion of what "byte" means in the C spec.
Yeah, it's one of those things that is not exactly intuitive if you didn't already know about it. The menagerie of other byte-ish-sized types added in recent versions of C and C++ probably doesn't help either.
Bit of a (not so?) fun fact: char* being a universal alias can lead to some potentially unexpected slowdowns [0], especially if the char* bit is behind a typedef.
It’s not any of that. It’s because of overcommit being the default for basically every Linux system. With that, malloc will never fail, and it’s the later access of that memory that will. In practice, you’ll virtually never see malloc actually return a failure, and so most software, no matter the language, is generally not robust to this condition.
recently noticed that iced (rust ui framework) crashes on oom - from the logs it seems to be aware and crash explicitly.
I wish it would just reduce the fps or hang a bit instead of crashing, perhaps use exponential backoff
How is it then that my stock KUbuntu system won't let me allocate 50,000,000,000 bytes?
$ uname -a
Linux boxcar 7.0.0-31-generic #31~24.04.1-Ubuntu SMP PREEMPT_DYNAMIC Mon Aug 10 09:38:02 UTC 2 x86_64 x86_64 x86_64 GNU/Linux
$ cat tmp.c
#include <stdio.h>
#include <stdlib.h>
int main() {
char *s = malloc(50000000000ULL);
if (s == NULL) {
printf("Boo, hoo!\n");
} else {
printf("Look at all that memory!\n");
}
return 0;
}
$ cc tmp.c
$ ./a.out
Boo, hoo!
As to the lack of robustness of most software, that's a fact. But, for example, Daniel Stenberg of curl fame is a developer of robust software who does not appreciate how Rust handles out-of-memory errors makes it impossible to implement libcurl in the way he expects a library to work. See https://www.youtube.com/watch?v=HFH2vZRTKrA&t=2080s from 4.5 years ago as an example.
To get back to the essay, I can easily understand why someone who has Rust as their first-and-only systems language, and therefore expects abort-when-out-of-memory, will not immediately consider how C allows a different approach to how to handle that condition, but instead will try to replicate Rust's behavior in C.
> How is it then that my stock KUbuntu system won't let me allocate 50,000,000,000 bytes?
I can't speak to the specific implementations, however servicing an allocation request has multiple fallible steps: reserve address space, then confirm/acquire/defer the physical memory backing the allocation.
My understanding is that overcommit allows deferring assigning physical memory to the allocation, however it could still fail to find a chunk of address space to service the allocation request.
For context, steveklabnik suggested that malloc will virtually never fail on basically every Linux system, because of overcommit being the default.
I showed a trivial example of malloc failing, when trying to allocate more space than on my machine.
Another example is when using resource limits. I've modified my code to malloc only 500M bytes, the set a virtual limit of 500000KiB, which works, then 40000KiB, which doesn't
$ grep malloc tmp.c
char *s = malloc(500000000ULL);
$ cc tmp.c
$ ulimit -S -v 500000
$ ./a.out
Look at all that memory!
$ ulimit -S -v 400000
$ ./a.out
Boo, hoo!
I clearly disagree with steveklabnik's because it's easy to demonstrate cases where malloc fails on Linux-based machines.
Not sure I agree with this pattern for strings, but if you're gonna multiply for an allocation like that and are concerned about consistent idioms, use calloc or [non-standard] mallocarray, which will check the multiplication for overflow.
I think it's because the new_uninit_slice call Rust will trigger a panic? Or abort? With little-to-no chance for recovery? (I know little about Rust.)
If so, I can see why someone that someone coming from Rust might consider exit() to be the appropriate solution for C, even for library code which should never be in charge of deciding how a program should exit.
I think the essay could be improved by highlighting the different worldviews.
I'm also old enough that
makes me nervous. Even for char -- I've never been on a system where sizeof(char) != 1 -- I want to see the sizeof included in the calculation, like: so I don't have to think about sizeof(char) being special.As long as I'm here, I'm a bit confused about the purpose of the "char* error_message" in the proposed Result. Why a char* vs a const char * or even better, an int with an error code? Who sees the message? Do we expect they know English, or will they be localized? Will the error message text be frozen forever, or might it change in the future?
And you never will, since sizeof(char) is guaranteed to always be 1.
I'm guessing you were thinking of CHAR_BIT != 8, but even then I'm not sure it would make a difference since malloc takes its argument size in bytes and a char more or less is a byte in C.
(Consider that char*s are also how you access the byte-level representation of objects in C. If chars were not the minimum addressable unit then that use wouldn't work)
https://smd.hu/Data/Analog/DSP/SHARC/C&C++%20Compiler%20&%20... says the cc21k compiler for ADSP-21xxx DSP systems has char as 32 bits signed, and that the compiler handles ANSI/ISO standard C.
So I don't believe your statement "sizeof(char) is guaranteed to always be 1" is correct.
> chars are also how you access the byte-level representation of objects in C
Where does the spec say that a char can be used to address any point in an object?
There's all sorts of oddities like tagged architectures which the C spec handles which I know essentially nothing about, but which break common expectations about how C works. I believe this is one of them.
I believe the following is undefined behavior in C, even though your compiler may let you do it, at least sometimes, and on modern desktop hardware:
I believe the following is the correct (or less incorrect) way to do it:Yes, but in C standardese a "byte" is not necessarily the 8 bits that it's normally thought to be these days. From C89 Section 2.2.4.2 Numerical Limits [-1]:
> maximum number of bits for smallest object that is not a bit-field (byte) CHAR_BIT 8
i.e., CHAR_BIT is the number of bits in a byte. C23 has a similar definition, and further defines CHAR_WIDTH that is defined to expand to the same value as CHAR_BIT.
> So I don't believe your statement "sizeof(char) is guaranteed to always be 1" is correct.
From C89 section 3.3.3.4 The sizeof operator [0]:
> When applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1.
This wording remains basically identical through C23 [1].
> Where does the spec say that a char* can be used to address any point in an object?
From C89 section 3.3 Expressions:
> An object shall have its stored value accessed only by an lvalue that has one of the following types:
> <snip>
> * a character type.
This also remains the case up through C23.
I think you're thinking of the strict aliasing rule with your example, but character types are one of the exceptions to said rule so I think your example is actually fully defined. It'd be UB if you casted to an incompatible type like a float, I believe.
[-1]: https://port70.net/%7Ensz/c/c89/c89-draft.html#2.2.4.2
[0]: https://port70.net/%7Ensz/c/c89/c89-draft.html#3.3.3.4
[1]: https://port70.net/%7Ensz/c/c23/n3220.html#6.5.4.4
[2]: https://port70.net/%7Ensz/c/c89/c89-draft.html#3.3
[3]: https://port70.net/%7Ensz/c/c23/n3220.html#6.5.1
So that scaling factor is placed into CHAR_BIT, and would be 64 on that DSP compiler.
> An object shall have its stored value accessed only by an lvalue that has one of the following types:
Yes, my confusion comes down to my confusion of what "byte" means in the C spec.
Thank you for your time in pointing this out.
Bit of a (not so?) fun fact: char* being a universal alias can lead to some potentially unexpected slowdowns [0], especially if the char* bit is behind a typedef.
[0]: https://travisdowns.github.io/blog/2019/08/26/vector-inc.htm...
It’s not any of that. It’s because of overcommit being the default for basically every Linux system. With that, malloc will never fail, and it’s the later access of that memory that will. In practice, you’ll virtually never see malloc actually return a failure, and so most software, no matter the language, is generally not robust to this condition.
To get back to the essay, I can easily understand why someone who has Rust as their first-and-only systems language, and therefore expects abort-when-out-of-memory, will not immediately consider how C allows a different approach to how to handle that condition, but instead will try to replicate Rust's behavior in C.
"I can program FORTRAN in any language." :)
I can't speak to the specific implementations, however servicing an allocation request has multiple fallible steps: reserve address space, then confirm/acquire/defer the physical memory backing the allocation.
My understanding is that overcommit allows deferring assigning physical memory to the allocation, however it could still fail to find a chunk of address space to service the allocation request.
For context, steveklabnik suggested that malloc will virtually never fail on basically every Linux system, because of overcommit being the default.
I showed a trivial example of malloc failing, when trying to allocate more space than on my machine.
Another example is when using resource limits. I've modified my code to malloc only 500M bytes, the set a virtual limit of 500000KiB, which works, then 40000KiB, which doesn't
I clearly disagree with steveklabnik's because it's easy to demonstrate cases where malloc fails on Linux-based machines.