前言
std::string大概是 C++ 里被用得最多、也被误解得最多的类型。常见的误解有三个:第一,以为它就是char数组的语法糖,本质上和char buf[100]差不多;第二,以为std::string一定在堆上分配内存,所以"性能肯定比栈上的字符数组差";第三,以为"模拟实现"就是把std::string的源码抄一遍——实际上标准只规定了接口的行为,内部怎么存、有没有短字符串优化(Small String Optimization, SSO)、扩容因子是多少,全是实现定义的。
本文分两半:前半讲std::string的接口和真正会用到的用法,后半写一个"能编译、能跑、行为正确"的简化版MyString,通过它把"拷贝控制"这件事讲清楚。最后给出几个真会踩的坑,尤其是c_str()的悬垂指针和迭代器失效。
目标读者是刚学完 C 字符串、准备用std::string替换strcpy/strcat的人。本文代码以 C++17 为基准,GCC 13 / Clang 17 / MSVC 19.3x 均可编译。
一、std::string 到底是什么
std::string不是标准库里的"一个类",而是一个类型别名:
// 概念示意,不是可直接编译的声明 namespace std { template<class CharT, class Traits = char_traits<CharT>, class Allocator = allocator<CharT>> class basic_string; using string = basic_string<char>; }也就是说,std::string是std::basic_string<char>。真正被标准规定的东西是basic_string的接口和复杂度要求;具体怎么实现由标准库厂商决定。你可以自己验证大小:
// C++17 #include <iostream> #include <string> #include <vector> int main() { std::cout << "sizeof(std::string) = " << sizeof(std::string) << '\n'; std::cout << "sizeof(std::vector<int>)= " << sizeof(std::vector<int>) << '\n'; std::string s = "hi"; std::cout << "size = " << s.size() << ", capacity = " << s.capacity() << '\n'; return 0; }这段代码在 GCC 13(libstdc++,默认的 C++11 ABI)上通常打印sizeof(std::string) = 32。为什么是 32 而不是"一个指针加两个整数"的 24?因为多数实现在对象内部留了一小块本地缓冲区做 SSO:短字符串直接存在对象里,不碰堆。
SSO 的阈值是彻头彻尾的实现细节,标准一个字都没规定。下面是三家主流实现的常见情况,仅供理解,具体数值请以你本地的sizeof和capacity()实测为准:
| 实现 | 所属编译器 | SSO 内部缓冲区常见容量 | sizeof(std::string)常见值 |
|---|---|---|---|
libstdc++(C++11 ABI,std::__cxx11::basic_string) | GCC 5+ | 15 个字符 + 结尾空字符 | 32 |
libstdc++(旧 ABI,_GLIBCXX_USE_CXX11_ABI=0) | GCC 5 之前 | 无 SSO(写时复制) | 8 |
| libc++ | Clang | 22 个字符 + 结尾空字符 | 24 |
| MSVC STL | MSVC | 15 个字符 + 结尾空字符 | 32 |
所以当有人问"std::string存 16 个字符会不会分配内存"时,正确答案是"取决于实现,GCC 的 libstdc++ 通常在 16 个字符时就已经转到堆上了,而 Clang 的 libc++ 要到 23 个字符才转"。别把它当标准断言。
二、真正会用到的接口
std::string的成员函数很多,但日常高频的其实就是下面这些。全部以标准的规定为准(需要精确签名时查 cppreference 或标准 [string] 一节)。
| 分类 | 成员函数 | 说明 |
|---|---|---|
| 容量 | size()/length() | 两者等价,返回字符个数(不含结尾空字符) |
| 容量 | capacity() | 当前已分配空间能容纳多少字符 |
| 容量 | empty() | 是否为空 |
| 容量 | reserve(n) | 预留至少 n 个字符的空间,避免多次扩容 |
| 容量 | resize(n) | 改变元素个数,多出的位置用char()填充 |
| 容量 | clear() | 清空内容,容量一般不变 |
| 访问 | operator[](i) | 不检查越界,越界是 UB |
| 访问 | at(i) | 越界抛std::out_of_range |
| 访问 | front()/back() | C++11 起;空串上调用是 UB |
| 访问 | data()/c_str() | 返回指向内部缓冲的指针;c_str()一定以'\0'结尾 |
| 修改 | push_back(c)/pop_back() | 追加/删除末尾字符 |
| 修改 | append(s)/operator+= | 追加;+=比append更常用也更短 |
| 修改 | insert(pos, s) | 在 pos 处插入 |
| 修改 | erase(pos, n) | 删除从 pos 起 n 个字符 |
| 查找 | find(s, pos = 0) | 返回首次出现位置,找不到返回npos |
| 查找 | rfind,find_first_of,find_first_not_of | 返回值同样用npos表示失败 |
| 截取 | substr(pos = 0, count = npos) | 返回新串(有拷贝),pos 越界抛std::out_of_range |
| 比较 | compare(s) | 返回负值/0/正值 |
关于data()有一个必须记住的版本差异:
| 标准版本 | const std::string上的data()返回 | 非 const 对象上的data()返回 |
|---|---|---|
| C++11 / C++14 | const char* | const char* |
| C++17 起 | const char* | char*(可写) |
也就是说,C++17 起非 const 的data()是可写的,但不能写超过size()的位置,也不能改写size()之后那个结尾空字符——那是 UB。
三、三个最常用的惯用法
用法一:拼接时先reserve。循环里反复+=会触发重新分配(reallocation):分配更大的块、把旧数据搬过去、释放旧块。搬一次就是 O(n)。如果提前知道大概长度,reserve能把这个开销压成一次。
// C++17 #include <string> #include <vector> #include <iostream> std::string join(const std::vector<std::string>& parts, char sep) { std::size_t total = 0; for (const auto& p : parts) total += p.size(); if (!parts.empty()) total += parts.size() - 1; std::string out; out.reserve(total); // 只预留一次 for (std::size_t i = 0; i < parts.size(); ++i) { if (i) out += sep; out += parts[i]; } return out; } int main() { std::vector<std::string> v{"alpha", "beta", "gamma"}; std::cout << join(v, ',') << '\n'; // alpha,beta,gamma return 0; }用法二:用find+substr切分,注意npos的比较方式。std::string::npos是static const size_type npos = -1(即size_type的最大值)。比较时要小心类型宽度:把find的结果存进int会在 64 位平台上截断。
// C++17 #include <string> #include <vector> std::vector<std::string> split(const std::string& s, char sep) { std::vector<std::string> out; std::string::size_type start = 0; while (true) { std::string::size_type pos = s.find(sep, start); if (pos == std::string::npos) { out.push_back(s.substr(start)); break; } out.push_back(s.substr(start, pos - start)); start = pos + 1; } return out; }用法三:需要 C 接口时用c_str(),但只在调用期间用。
// C++17 #include <string> #include <cstdio> int main() { std::string name = "cpp"; // 直接把指针交给 C 函数,printf 在本次调用内使用它,安全 std::printf("%s\n", name.c_str()); return 0; }c_str()返回的指针在任何会修改这个 string 的操作之后都可能失效(包括push_back、+=、reserve导致的扩容)。这不是"可能失效"的模糊说法——标准规定这类操作会使指向元素的指针/引用失效的规则适用于data(),c_str()同理。
实战:一个可编译的简化版 MyString
下面这个类不追求接口完备,只实现最核心的部分:构造、析构、拷贝构造、拷贝赋值、移动构造、移动赋值、size、c_str、operator[]、operator+=,并且不实现 SSO(所有非空数据都在堆上)。通过它可以看清"拷贝控制"的五件事。
// C++17,单文件可直接编译:g++ -std=c++17 -Wall -Wextra mystring.cpp #include <cstddef> #include <cstring> #include <iostream> #include <utility> class MyString { public: // 默认构造:空串也要有一个合法的 buf_,保证 c_str() 可用 MyString() : size_(0), cap_(0), buf_(new char[1]) { buf_[0] = '\0'; } MyString(const char* s) { size_ = std::strlen(s); cap_ = size_; buf_ = new char[size_ + 1]; std::memcpy(buf_, s, size_ + 1); // 连结尾 '\0' 一起拷 } // 拷贝构造:深拷贝 MyString(const MyString& other) { size_ = other.size_; cap_ = other.size_; buf_ = new char[size_ + 1]; std::memcpy(buf_, other.buf_, size_ + 1); } // 移动构造:接管别人的缓冲区,并把对方置为有效但空的状态 MyString(MyString&& other) noexcept : size_(other.size_), cap_(other.cap_), buf_(other.buf_) { other.size_ = 0; other.cap_ = 0; other.buf_ = new char[1]; // 对方仍然要能析构、能 c_str() other.buf_[0] = '\0'; } MyString& operator=(const MyString& other) { if (this != &other) { // 必须自赋值检查,否则自己释放自己 MyString tmp(other); // 先拷贝,异常安全 swap(tmp); } return *this; } MyString& operator=(MyString&& other) noexcept { if (this != &other) { MyString tmp(std::move(other)); swap(tmp); } return *this; } ~MyString() { delete[] buf_; } void swap(MyString& other) noexcept { std::swap(size_, other.size_); std::swap(cap_, other.cap_); std::swap(buf_, other.buf_); } std::size_t size() const { return size_; } bool empty() const { return size_ == 0; } const char* c_str() const { return buf_; } char& operator[](std::size_t i) { return buf_[i]; } // 不检查越界 const char& operator[](std::size_t i) const { return buf_[i]; } void reserve(std::size_t n) { if (n <= cap_) return; char* nb = new char[n + 1]; std::memcpy(nb, buf_, size_ + 1); delete[] buf_; buf_ = nb; cap_ = n; } MyString& operator+=(const MyString& rhs) { if (rhs.size_ == 0) return *this; if (size_ + rhs.size_ > cap_) { reserve(size_ + rhs.size_); // 简化策略:需要多少要多少 } std::memcpy(buf_ + size_, rhs.buf_, rhs.size_ + 1); size_ += rhs.size_; return *this; } MyString& operator+=(const char* rhs) { return *this += MyString(rhs); } private: std::size_t size_; std::size_t cap_; char* buf_; }; MyString operator+(MyString lhs, const MyString& rhs) { lhs += rhs; // 传值 + 返回,天然享受移动语义 return lhs; } std::ostream& operator<<(std::ostream& os, const MyString& s) { return os << s.c_str(); } int main() { MyString a("Hello"); MyString b = a; // 拷贝构造 b += MyString(", world"); std::cout << b << " (size=" << b.size() << ")\n"; MyString c = a + b; // 移动构造 std::cout << c << " (size=" << c.size() << ")\n"; MyString d; d = std::move(c); // 移动赋值 std::cout << d << '\n'; a = a; // 自赋值,必须是安全的 std::cout << a << '\n'; return 0; }这段代码在 GCC 13 / Clang 17 上用-std=c++17 -Wall -Wextra编译应无警告。几个设计点值得说明:
- 默认构造也分配 1 字节,这样
c_str()永远返回一个合法的、以'\0'结尾的指针。 - 移动构造里
noexcept不是装饰。标准容器(比如std::vector)在扩容时判断元素的移动构造函数是否noexcept:是则用移动,否则为了强异常保证会退回到拷贝。标上noexcept能让容器选择更省的路径。 - 移动后源对象仍处于"有效但未指定"状态。我在移动构造里给源对象重新分配了空缓冲区,这比"留一个空指针"更安全(后者会让源对象的
c_str()直接崩溃)。 - 赋值用"拷贝并交换"(copy-and-swap),天然处理自赋值,且异常安全。
operator+按值接收左操作数,于是lhs本身就是一份拷贝,可以直接在上面追加再返回——返回时享受移动语义。
需要坦白一点:这个简化版的移动构造里做了一次new char[1],却把移动构造标成了noexcept。一旦这次分配失败抛出std::bad_alloc,程序会直接调用std::terminate。真实的标准库实现靠 SSO 或共享的空串静态对象避免这次分配,它们的移动构造是真的不会失败。写生产代码时不要照抄这一点:要么让移动构造真的不做可能抛异常的事,要么就别标noexcept——而不标noexcept又会失去容器扩容时优先移动的机会,这正是 SSO 重要性的来源之一。
常见坑点
坑 1:把c_str()的返回值存下来跨语句使用。
❌
const char* p = s.c_str(); s += "x"; // 可能触发扩容,p 变成悬垂指针 std::printf("%s\n", p); // UB,标准不保证任何行为✅ 只在同一个表达式/同一个调用内用c_str();确实要留住,就复制一份到std::string或std::vector<char>。
坑 2:substr越界不会报错,但at会。
❌std::string s = "abc"; char c = s[10];——operator[]不检查,越界是 UB。
✅ 用s.at(10),它在越界时抛std::out_of_range。注意substr(pos)在pos > size()时也抛std::out_of_range。
坑 3:把npos塞进int。
❌
int pos = s.find("x"); // npos 在 64 位平台被截断,判断结果错乱 if (pos == std::string::npos) // 类型不匹配,比较结果不可靠✅ 用std::string::size_type(或auto)接收find的返回值,再和std::string::npos比较。
坑 4:以为reserve之后指针就永远稳定了。
❌ 在循环里reserve一次,然后一路追加并缓存data()返回的指针。
✅reserve只是减少扩容次数;任何可能改动size或capacity的操作都可能让之前的指针失效。要长期持有就存下标,不存指针。
坑 5:中文字符串用size()当"字数"。
❌std::string s = "中文";然后认为s.size() == 2。UTF-8 下一个汉字通常是 3 字节,因此s.size()是 6。
✅ 明确区分"字节数"和"字符数"。字符级处理需要宽字符、std::u8string(C++20)或第三方 Unicode 库;只想知道 UTF-8 码点个数,可以手写一个按首字节判断续字节个数的计数器。
坑 6:循环里s = s + "x";而不是s += "x";。
❌
for (int i = 0; i < 1000; ++i) s = s + "x"; // 每次都构造临时串再拷贝赋值✅s += "x";——operator+=直接在原串上追加,不需要构造完整副本。
总结
| 主题 | 结论 |
|---|---|
std::string的本质 | std::basic_string<char>的别名,接口由标准规定,布局由实现决定 |
| SSO | 实现细节,GCC 的 libstdc++ 与 MSVC STL 常见 15 字符,Clang 的 libc++ 常见 22 字符 |
| 取字符指针 | 用c_str(),且只在调用期间使用;C++17 起非 const 的data()可写 |
| 查找失败 | 返回std::string::npos,必须用std::string::size_type接收 |
| 效率 | 循环拼接前先reserve;用+=而不是s = s + ... |
| 实现自己的字符串类 | 必须成套实现拷贝构造/拷贝赋值/移动构造/移动赋值/析构,移动构造标noexcept |
std::string的接口看着平易近人,真正的难点在两个地方:一是生命周期——所有返回指针的函数(c_str、data)都把"什么时候会失效"的责任交给了调用者;二是实现差异——SSO、容量增长策略这些看起来"应该是标准"的东西其实全是厂商自由发挥。把这两点记牢,用std::string就很少会出问题;自己写一个简化版,则是把拷贝控制这套规则真正内化的最快办法。